VLDB 2026 Research / reviewers in the wild / expert
Junwen Wang
dblp:29/5437
· DBLP profile ↗
38ranked-venue papers
9as first author
17since 2021 · last 2026
0000-0002-4432-4707ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 31 · 6 first-author · 13 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021Computer networks · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Partner or subordinate? Heterogeneous effects of crane operator workload and collision risk under a binary HMI paradigm
Zhijiang Wu, Yaru Zhu, Junwen Wang |
Adv. Eng. Informatics | 3 |
| 2026 | ceQTL: a co-expression QTL model to detect a variant that affects transcription factor binding and its target regulationabstractExpression quantitative trait locus (eQTL) mapping is used to identify a functional link between a genomic variant, such as single nucleotide polymorphism (SNP), and gene expression (often close-by pair for cis-eQTL) by linear regression, commonly done when matching genotype and expression data from same individuals are available. Millions of significant eQTLs have been reported by both individual studies and coordinated large consortiums such as Genotype-Tissue Expression project (GTEx). A significant eQTL association does not establish a causal relationship or provide any underlying mechanism so further investigation is needed to understand how a SNP impacts gene expression. One of the plausible explanations for eQTL is that a genomic variant affects transcription factor (TF) binding and thus impacts its regulation on its target genes (TGs). However, data-driven or formal statistical methods to prove that hypothesis are still lacking. To address the gap, we propose a new method called differential co-expression QTL (ceQTL) among different alleles using Chow statistics to specifically detect eQTLs that are modulated by a particular TF. We start with building a trio of TF, its TG, and related SNP and then test the significant coefficient difference among different levels of SNP in terms of TF and TG correlation. We applied this ceQTL model to simulated data and the lung tissue datasets from the GTEx project. The simulated data results showed that the model was robust to detect true ceQTLs at variable sample sizes and different minor allele frequencies as measured by Area Under the Curve (AUC). In normal lung tissue, a small fraction of eQTLs were found to have strong ceQTLs, i.e., eQTLs where SNP affects gene expression though TF binding. Some ceQTLs may not be detected by traditional eQTL analysis. Our tool also performed a TF binding affinity analysis to add another layer of evidence for functional interpretation. Comparisons with other similar tools were also presented. In summary, ceQTL analysis provides a more interpretable and biological insight into the mechanism of eQTL, which would help us better understand how genomic variants affect phenotypes and diseases. Panwen Wang, Yanxi Chen 0002, Li Liu 0035, Junwen Wang, Zhifu Sun |
Briefings Bioinform. | 6 |
| 2026 | DeepPTMPred: a multi-modal deep learning framework for accurate prediction of protein post-translational modification sitesabstractPost-translational modifications are important for regulating cellular functions. Although traditional experimental methods accurately identify PTM sites, they are time-consuming. In this study, we propose a novel model capable of predicting 17 types of PTMs through multi-modal integration and AlphaFold predictions. Our model employs an enhanced CNN-transformer architecture to capture local dependencies within the sequence, while incorporating structural features and evolutionary patterns to effectively capture complex spatial relationships and global contextual dependencies. Through rigorous cross-validation and testing, our model demonstrates exceptional performance, achieving area under the curve scores of 96.5%, 91.6%, 91.0%, and 89.5% for the prediction of hydroxylation, malonylation, O-linked glycosylation, and phosphorylation, respectively. Notably, our model accurately identified known phosphorylation sites on tau and two recently identified residues linked to pre-tangle stages and early Alzheimer's disease pathology. This work not only deepens the understanding of PTMs but also holds promise for advancing future research in the prediction of PTM sites and functional annotation. Chenkui Wang, Qianhui Jiang, Jiahui Guan, Zimeng Chen, Xiaoling Lu, Jing Qin 0004, Junwen Wang |
Briefings Bioinform. | 10 |
| 2026 | OOD-SEG: Exploiting out-of-distribution detection techniques for learning image segmentation from sparse multi-class positive-only annotationsabstract• We propose a segmentation framework for learning from sparse positive-only multi-class annotations without background labels. • We model implicit background via pixel-wise OOD detection, enabling separation of in-distribution data from negative/unknown regions. • We propose a two-level cross-validation strategy that addressing the scarcity of OOD datasets and the lack of established evaluation metrics in segmentation. • We introduce a novel convolutional adaptation of an established OOD detection method originally designed for classification purposes. • We demonstrate robustness and generalisation through experiments on multi-class hyperspectral and RGB surgical imaging datasets. Despite significant advancements, segmentation based on deep neural networks in medical and surgical imaging faces several challenges, two of which we aim to address in this work. First, acquiring complete pixel-level segmentation labels for medical images is time-consuming and requires domain expertise. Second, typical segmentation pipelines cannot detect out-of-distribution (OOD) pixels, leaving them prone to spurious outputs during deployment. In this work, we propose a novel segmentation approach which broadly falls within the positive-unlabelled (PU) learning paradigm and exploits tools from OOD detection techniques. Our framework learns only from sparsely annotated pixels from multiple positive-only classes and does not use any annotation for the background class. These multi-class positive annotations naturally fall within the in-distribution (ID) set. Unlabelled pixels may contain positive classes but also negative ones, including what is typically referred to as background in standard segmentation formulations. To the best of our knowledge, this work is the first to formulate multi-class segmentation with sparse positive-only annotations as a pixel-wise PU learning problem and to address it using OOD detection techniques. Here, we forgo the need for background annotation and consider these together with any other unseen classes as part of the OOD set. Our framework can integrate, at a pixel-level, any OOD detection approaches designed for classification tasks. To address the lack of existing OOD datasets and established evaluation metric for medical image segmentation, we propose a cross-validation strategy that treats held-out labelled classes as OOD. Extensive experiments on both multi-class hyperspectral and RGB surgical imaging datasets demonstrate the robustness and generalisation capability of our proposed framework. Junwen Wang, Oscar MacCormac, Jonathan Shapey, Tom Vercauteren |
Medical Image Anal. | 1 |
| 2025 | Tree-Based Semantic Losses: Application to Sparsely-Supervised Large Multi-class Hyperspectral Segmentation
Junwen Wang, Oscar MacCormac, William Rochford, Aaron Kujawa, Jonathan Shapey, Tom Vercauteren |
MICCAI (8) | 1 |
| 2025 | MSF-CPMP: a novel multi-source feature fusion model for prediction of cyclic peptide membrane permeabilityabstractAbstract Motivation Membrane permeability represents a critical bottleneck in cyclic peptide drug development, limiting the clinical translation of these therapeutically attractive molecules despite their inherent stability and structural diversity. Current computational models for predicting cyclic peptide membrane permeability (CPMP) exhibit insufficient accuracy for early-stage drug screening, hampering the efficiency of lead optimization in pharmaceutical pipelines. Methodology In this study, we introduce a novel multi-source feature fusion model called MSF-CPMP, which aims to increase the accuracy of predicted CPMP. The MSF-CPMP model incorporates three features extracted from SMILES sequences, graph-based molecular structures, and physicochemical properties of cyclic peptides. Results By benchmarking with other machine learning and deep learning-based methods, MSF-CPMP achieved the highest levels of the evaluation metrics such as accuracy of 0.9062 and AUROC of 0.9546, and further validated MSF-CPMP robustness in learning capabilities and efficacy of its multi-source fusion. MSF-CPMP has been validated on FDA-approved cyclic peptide therapeutics, demonstrating strong predictive power for clinical candidates. Our result demonstrates that MSF-CPMP outperforms other methods in predicting CPMP, providing a practical computational tool for accelerating drug screening and reducing attrition rates in cyclic peptide drug discovery, thereby advancing precision-guided pharmaceutical development. Availability Code is available at https://github.com/wanglabhku/MSF-CPMP Zimeng Chen, Zhuxuan Wan, Qianhui Jiang, Xiaoling Lu, Jing Qin 0004, Junwen Wang |
Briefings Bioinform. | 9 |
| 2025 | Graph-RPI: predicting RNA-protein interactions via graph autoencoder and self-supervised learning strategiesabstractRNA-protein interactions (RPIs) are essential for many biological functions and are associated with various diseases. Traditional methods for detecting RPIs are labor-intensive and costly, necessitating efficient computational methods. In this study, we proposed a novel sequence-based RPI prediction framework based on graph neural networks (GNNs) that addressed key limitations of existing methods, such as inadequate feature integration and negative sample construction. Our method represented RNAs and proteins as nodes in a unified interaction graph, enhancing the representation of RPI pairs through multi-feature fusion and employing self-supervised learning strategies for model training. The model's performance was validated through five-fold cross-validation, achieving accuracy of 0.880, 0.811, 0.950, 0.979, 0.910, and 0.924 on the RPI488, RPI369, RPI2241, RPI1807, RPI1446, and RPImerged datasets, respectively. Additionally, in cross-species generalization tests, our method outperformed existing methods, achieving an overall accuracy of 0.989 across 10 093 RPI pairs. Compared with other state-of-the-art RPI prediction methods, our approach demonstrates greater robustness and stability in RPI prediction, highlighting its potential for broad biological applications and large-scale RPI analysis. Jiahui Guan, Lantian Yao, Peilin Xie, Dian Meng, Tzong-Yi Lee, Junwen Wang, Ying-Chih Chiang |
Briefings Bioinform. | 7 |
| 2025 | A self-attention-driven deep learning framework for inference of transcriptional gene regulatory networksabstractThe interactions between transcription factors (TFs) and the target genes could provide a basis for constructing gene regulatory networks (GRNs) for mechanistic understanding of various biological complex processes. From gene expression data, particularly single-cell transcriptomic data containing rich cell-to-cell variations, it is highly desirable to infer TF-gene interactions (TGIs) using deep learning technologies. Numerous models or software including deep learning-based algorithms have been designed to identify transcriptional regulatory relationships between TFs and the downstream genes. However, these methods do not significantly improve predictions of TGIs due to some limitations regarding constructing underlying interactive structures linking regulatory components. In this study, we introduce a deep learning framework, DeepTGI, that encodes gene expression profiles from single-cell and/or bulk transcriptomic data and predicts TGIs with high accuracy. Our approach could fuse the features extracted from Auto-encoder with self-attention mechanism and other networks and could transform multihead attention modules to define representative features. By comparing it with other models or methods, DeepTGI exhibits its superiority to identify more potential TGIs and to reconstruct the GRNs and, therefore, could provide broader perspectives for discovery of more biological meaningful TGIs and for understanding transcriptional gene regulatory mechanisms. Le Zhong, Zhuobin Chen, Yanjia Yu, Jing Qin 0004, Junwen Wang |
Briefings Bioinform. | 8 |
| 2025 | TPepPro: a deep learning model for predicting peptide-protein interactionsabstractMOTIVATION: Peptides and their derivatives hold potential as therapeutic agents. The rising interest in developing peptide drugs is evidenced by increasing approval rates by the FDA of USA. To identify the most potential peptides, study on peptide-protein interactions (PepPIs) presents a very important approach but poses considerable technical challenges. In experimental aspects, the transient nature of PepPIs and the high flexibility of peptides contribute to elevated costs and inefficiency. Traditional docking and molecular dynamics simulation methods require substantial computational resources, and the predictive accuracy of their results remain unsatisfactory. RESULTS: To address this gap, we proposed TPepPro, a Transformer-based model for PepPI prediction. We trained TPepPro on a dataset of 19,187 pairs of peptide-protein complexes with both sequential and structural features. TPepPro utilizes a strategy that combines local protein sequence feature extraction with global protein structure feature extraction. Moreover, TPepPro optimizes the architecture of structural featuring neural network in BN-ReLU arrangement, which notably reduced the amount of computing resources required for PepPIs prediction. According to comparison analysis, the accuracy reached 0.855 in TPepPro, achieving an 8.1% improvement compared to the second-best model TAGPPI. TPepPro achieved an AUC of 0.922, surpassing the second-best model TAGPPI with 0.844. Moreover, the newly developed TPepPro identify certain PepPIs that can be validated according to previous experimental evidence, thus indicating the efficiency of TPepPro to detect high potential PepPIs that would be helpful for amino acid drug applications. AVAILABILITY AND IMPLEMENTATION: The source code of TPepPro is available at https://github.com/wanglabhku/TPepPro. Xiaohong Jin, Zimeng Chen, Qianhui Jiang, Zhuobin Chen, Jing Qin 0004, Junwen Wang |
Bioinform. | 9 |
| 2025 | scDILT: A Model-Based and Constrained Deep Learning Framework for Single-Cell Data Integration, Label Transferring, and ClusteringabstractThe scRNA-seq technology enables high-resolution profiling and analysis of individual cells. The increasing availability of datasets and advancements in technology have prompted researchers to integrate existing annotated datasets with newly sequenced datasets for a more comprehensive analysis. It is important to ensure that the integration of new datasets does not alter the cell clusters defined in the old/reference datasets. Although several methods have been developed for scRNA-seq data integration, there is currently a lack of tools that can simultaneously achieve the aforementioned objectives. Therefore, in this study, we have introduced a novel tool called scDILT, which leverages a conditional autoencoder and deep embedding clustering to effectively remove batch effects among different datasets. Moreover, scDILT utilizes homogeneous constraints to preserve the cell-type/clustering patterns observed in the reference datasets, while employing heterogeneous constraints to map cells in the new datasets to the annotated cell clusters in the reference datasets. We have conducted extensive experiments to demonstrate that scDILT outperforms other methods in terms of data integration, as confirmed by evaluations on simulated and real datasets. Furthermore, we have shown that scDILT can be successfully applied to integrate multi-omics single-cell datasets. Based on these findings, we conclude that scDILT holds great promise as a tool for integrating single-cell datasets derived from different batches, experiments, times, or interventions. Jianlan Ren, Junwen Wang, Zhi Wei 0001 |
IEEE Trans. Comput. Biol. Bioinform. | 4 |
| 2025 | CapUBS: Capsule Network-Based Band Selection for Underwater Hyperspectral ImageryabstractUnderwater hyperspectral imagery (HSI) finds extensive applications in critical domains such as marine target detection. However, spectral offset caused by underwater light propagation presents significant challenges for existing hyperspectral band selection methods, which often fail to capture inter-band correlations in underwater environments effectively. To address this issue, this paper proposes a novel underwater hyperspectral band selection method based on capsule network (CapUBS) designed to mitigate spectral offset variations and reduce redundancy in band selection for underwater HSI. Especially, the dual-stream latent variable encoding module takes the original and averaged spectra as inputs and employs deep feature extraction blocks to capture local details and global semantic features, thereby alleviating spectral variability. The capsule network-based band characterization module utilizes a dynamic routing mechanism to generate digital capsules representing joint multiband features. This module characterizes the multi-dimensional spectral attributes via vector directions and quantifies the relative importance of spectral bands using vector norms. The feature separation factor evaluation module computes the initial band importance ranking to refine the selection process further. It evaluates a feature separation factor to identify key bands with minimal redundancy and high representativeness, resulting in an optimal subset of spectral bands. Finally, the key band feature reconstruction module leverages these selected bands to reconstruct the full-band spectrum accurately. Comprehensive quantitative and qualitative analyses conducted on four classical hyperspectral datasets and underwater hyperspectral experimental datasets demonstrate the superior performance of the proposed method compared to state-of-the-art techniques. The code will be available at: https://github.com/lq132069/CapUBS. Qi Li 0054, Junwen Wang, Xingyuan Zu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MultiSC: a deep learning pipeline for analyzing multiomics single-cell dataabstractSingle-cell technologies enable researchers to investigate cell functions at an individual cell level and study cellular processes with higher resolution. Several multi-omics single-cell sequencing techniques have been developed to explore various aspects of cellular behavior. Using NEAT-seq as an example, this method simultaneously obtains three kinds of omics data for each cell: gene expression, chromatin accessibility, and protein expression of transcription factors (TFs). Consequently, NEAT-seq offers a more comprehensive understanding of cellular activities in multiple modalities. However, there is a lack of tools available for effectively integrating the three types of omics data. To address this gap, we propose a novel pipeline called MultiSC for the analysis of MULTIomic Single-Cell data. Our pipeline leverages a multimodal constraint autoencoder (single-cell hierarchical constraint autoencoder) to integrate the multi-omics data during the clustering process and a matrix factorization-based model (scMF) to predict target genes regulated by a TF. Moreover, we utilize multivariate linear regression models to predict gene regulatory networks from the multi-omics data. Additional functionalities, including differential expression, mediation analysis, and causal inference, are also incorporated into the MultiSC pipeline. Extensive experiments were conducted to evaluate the performance of MultiSC. The results demonstrate that our pipeline enables researchers to gain a comprehensive view of cell activities and gene regulatory networks by fully leveraging the potential of multiomics single-cell data. By employing MultiSC, researchers can effectively integrate and analyze diverse omics data types, enhancing their understanding of cellular processes. Siqi Jiang, Zhi Wei 0001, Junwen Wang |
Briefings Bioinform. | 5 |
| 2023 | The prognostic value and immune landscaps of m6A/m5C-related lncRNAs signature in the low grade gliomaabstractBACKGROUND: N6-methyladenosine (m6A) and 5-methylcytosine (m5C) are the main RNA methylation modifications involved in the oncogenesis of cancer. However, it remains obscure whether m6A/m5C-related long non-coding RNAs (lncRNAs) affect the development and progression of low grade gliomas (LGG). METHODS: We summarized 926 LGG tumor samples with RNA-seq data and clinical information from The Cancer Genome Atlas and Chinese Glioma Genome Atlas. 105 normal brain samples with RNA-seq data from the Genotype Tissue Expression project were collected for control. We obtained a molecular classification cluster from the expression pattern of sreened lncRNAs. The least absolute shrinkage and selection operator Cox regression was employed to construct a m6A/m5C-related lncRNAs prognostic signature of LGG. In vitro experiments were employed to validate the biological functions of lncRNAs in our risk model. RESULTS: The expression pattern of 14 sreened highly correlated lncRNAs could cluster samples into two groups, in which various clinicopathological features and the tumor immune microenvironment were significantly distinct. The survival time of cluster 1 was significantly reduced compared with cluster 2. This prognostic signature is based on 8 m6A/m5C-related lncRNAs (GDNF-AS1, HOXA-AS3, LINC00346, LINC00664, LINC00665, MIR155HG, NEAT1, RHPN1-AS1). Patients in the high-risk group harbored shorter survival times. Immunity microenvironment analysis showed B cells, CD4 + T cells, macrophages, and myeloid-derived DC cells were significantly increased in the high-risk group. Patients in high-risk group had the worse overall survival time regardless of followed TMZ therapy or radiotherapy. All observed results from the TCGA-LGG cohort could be validated in CGGA cohort. Afterwards, LINC00664 was found to promote cell viability, invasion and migration ability of glioma cells in vitro. CONCLUSION: Our study elucidated a prognostic prediction model of LGG by 8 m6A/m5C methylated lncRNAs and a critical lncRNA regulation function involved in LGG progression. High-risk patients have shorter survival times and a pro-tumor immune microenvironment. Chaoxi Li, Yiwei Qi, Junwen Wang, Chao You, Haohao Huang |
BMC Bioinform. | 6 |
| 2023 | RFIA-Net: Rich CNN-transformer network based on asymmetric fusion feature aggregation to classify stage I multimodality oesophageal cancer images
Zhicheng Zhou 0001, Long Yu 0001, Shengwei Tian, Guangli Xiao, Junwen Wang, Shaofeng Zhou |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | An angular shrinkage BERT model for few-shot relation extraction with none-of-the-above detection
Junwen Wang, Yongbin Gao, Zhijun Fang 0001 |
Pattern Recognit. Lett. | 1 |
| 2022 | CONST: Exploiting Spatial-Temporal Correlation for Multi-Gateway based Reliable LoRa ReceptionabstractAs a representative technology of low power wide area network, LoRa has been widely adopted to many applications. A fundamental question in LoRa is how to improve its reception quality in ultra-low SNR scenarios. Different from existing studies that exploit either spatial or temporal correlation for LoRa reception recovery, this paper jointly leverages the fine-grained spatial-temporal correlation among multiple gateways. We exploit the spatial and temporal correlation in LoRa packets to jointly process received signals so that the fine-grained offsets including Central Frequency Offset (CFO), Sampling Time Offset (STO) and Sampling Frequency Offset (SFO) are well compensated, and signals from multiple gateways are combined coherently. Moreover, a deep learning based soft decoding scheme is developed to integrate the energy distribution of each symbol into the decoder to further enhance the coding gain in a LoRa packet. We evaluate our work with commodity LoRa devices (i.e., Semtech SX1278) and gateways (i.e., USRP-B210) in both indoor and outdoor environments. Extensive experiment results show that our work achieves 4.6dB higher signal-to-noise ratio (SNR) and 1.5× lower bit error rate (BER) compared with existing approaches. Weiwei Chen 0004, Junwen Wang, Shuai Wang 0008, Tian He 0001 |
ICNP | 3 |
| 2021 | Cell fate conversion prediction by group sparse optimization method utilizing single-cell and bulk OMICs dataabstractCell fate conversion by overexpressing defined factors is a powerful tool in regenerative medicine. However, identifying key factors for cell fate conversion requires laborious experimental efforts; thus, many of such conversions have not been achieved yet. Nevertheless, cell fate conversions found in many published studies were incomplete as the expression of important gene sets could not be manipulated thoroughly. Therefore, the identification of master transcription factors for complete and efficient conversion is crucial to render this technology more applicable clinically. In the past decade, systematic analyses on various single-cell and bulk OMICs data have uncovered numerous gene regulatory mechanisms, and made it possible to predict master gene regulators during cell fate conversion. By virtue of the sparse structure of master transcription factors and the group structure of their simultaneous regulatory effects on the cell fate conversion process, this study introduces a novel computational method predicting master transcription factors based on group sparse optimization technique integrating data from multi-OMICs levels, which can be applicable to both single-cell and bulk OMICs data with a high tolerance of data sparsity. When it is compared with current prediction methods by cross-referencing published and validated master transcription factors, it possesses superior performance. In short, this method facilitates fast identification of key regulators, give raise to the possibility of higher successful conversion rate and in the hope of reducing experimental cost. Jing Qin 0004, Jen-Chih Yao, Ricky Wai Tak Leung, Yongqiang Zhou, Junwen Wang |
Briefings Bioinform. | 7 |
| 2020 | Methods and resources to access mutation-dependent effects on cancer drug treatmentabstractIn clinical cancer treatment, genomic alterations would often affect the response of patients to anticancer drugs. Studies have shown that molecular features of tumors could be biomarkers predictive of sensitivity or resistance to anticancer agents, but the identification of actionable mutations are often constrained by the incomplete understanding of cancer genomes. Recent progresses of next-generation sequencing technology greatly facilitate the extensive molecular characterization of tumors and promote precision medicine in cancers. More and more clinical studies, cancer cell lines studies, CRISPR screening studies as well as patient-derived model studies were performed to identify potential actionable mutations predictive of drug response, which provide rich resources of molecularly and pharmacologically profiled cancer samples at different levels. Such abundance of data also enables the development of various computational models and algorithms to solve the problem of drug sensitivity prediction, biomarker identification and in silico drug prioritization by the integration of multiomics data. Here, we review the recent development of methods and resources that identifies mutation-dependent effects for cancer treatment in clinical studies, functional genomics studies and computational studies and discuss the remaining gaps and future directions in this area. Hongcheng Yao, Xinyi Qian, Junwen Wang, Pak Chung Sham, Mulin Jun Li |
Briefings Bioinform. | 4 |
| 2020 | Somatic selection distinguishes oncogenes and tumor suppressor genesabstractMOTIVATION: Functions of cancer driver genes vary substantially across tissues and organs. Distinguishing passenger genes, oncogenes (OGs) and tumor-suppressor genes (TSGs) for each cancer type is critical for understanding tumor biology and identifying clinically actionable targets. Although many computational tools are available to predict putative cancer driver genes, resources for context-aware classifications of OGs and TSGs are limited. RESULTS: We show that the direction and magnitude of somatic selection of protein-coding mutations are significantly different for passenger genes, OGs and TSGs. Based on these patterns, we develop a new method (genes under selection in tumors) to discover OGs and TSGs in a cancer-type specific manner. Genes under selection in tumors shows a high accuracy (92%) when evaluated via strict cross-validations. Its application to 10 172 tumor exomes found known and novel cancer drivers with high tissue-specificities. In 11 out of 13 OGs shared among multiple cancer types, we found functional domains selectively engaged in different cancers, suggesting differences in disease mechanisms. AVAILABILITY AND IMPLEMENTATION: An R implementation of the GUST algorithm is available at https://github.com/liliulab/gust. A database with pre-computed results is available at https://liliulab.shinyapps.io/gust. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pramod Chandrashekar, Navid Ahmadinejad, Junwen Wang, Aleksandar Sekulic, Jan B. Egan, Yan W. Asmann, Sudhir Kumar 0001, Carlo Maley, Li Liu 0035 |
Bioinform. | 3 |
| 2019 | Evaluation of tools for highly variable gene discovery from single-cell RNA-seq dataabstractTraditional RNA sequencing (RNA-seq) allows the detection of gene expression variations between two or more cell populations through differentially expressed gene (DEG) analysis. However, genes that contribute to cell-to-cell differences are not discoverable with RNA-seq because RNA-seq samples are obtained from a mixture of cells. Single-cell RNA-seq (scRNA-seq) allows the detection of gene expression in each cell. With scRNA-seq, highly variable gene (HVG) discovery allows the detection of genes that contribute strongly to cell-to-cell variation within a homogeneous cell population, such as a population of embryonic stem cells. This analysis is implemented in many software packages. In this study, we compare seven HVG methods from six software packages, including BASiCS, Brennecke, scLVM, scran, scVEGs and Seurat. Our results demonstrate that reproducibility in HVG analysis requires a larger sample size than DEG analysis. Discrepancies between methods and potential issues in these tools are discussed and recommendations are made. Shun H. Yip, Pak Chung Sham, Junwen Wang |
Briefings Bioinform. | 3 |
| 2019 | MetaMarker: a pipeline for de novo discovery of novel metagenomic biomarkersabstractSUMMARY: We present MetaMarker, a pipeline for discovering metagenomic biomarkers from whole-metagenome sequencing samples. Different from existing methods, MetaMarker is based on a de novo approach that does not require mapping raw reads to a reference database. We applied MetaMarker on whole-metagenome sequencing of colorectal cancer (CRC) stool samples from France to discover CRC specific metagenomic biomarkers. We showed robustness of the discovered biomarkers by validating in independent samples from Hong Kong, Austria, Germany and Denmark. We further demonstrated these biomarkers could be used to build a machine learning classifier for CRC prediction. AVAILABILITY AND IMPLEMENTATION: MetaMarker is freely available at https://bitbucket.org/mkoohim/metamarker under GPLv3 license. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mohamad Koohi-Moghadam, Mitesh J. Borad, Nhan L. Tran, Kristin R. Swanson, Lisa Boardman, Hong-Zhe Su, Junwen Wang |
Bioinform. | 7 |
| 2018 | Inferring RNA sequence preferences for poorly studied RNA-binding proteins based on co-evolutionabstractBACKGROUND: Characterizing the binding preference of RNA-binding proteins (RBP) is essential for us to understand the interaction between an RBP and its RNA targets, and to decipher the mechanism of post-transcriptional regulation. Experimental methods have been used to generate protein-RNA binding data for a number of RBPs in vivo and in vitro. Utilizing the binding data, a couple of computational methods have been developed to detect the RNA sequence or structure preferences of the RBPs. However, the majority of RBPs have not yet been experimentally characterized and lack RNA binding data. For these poorly studied RBPs, the identification of their binding preferences cannot be performed by most existing computational methods because the experimental binding data are prerequisite to these methods. RESULTS: Here we propose a new method based on co-evolution to predict the sequence preferences for the poorly studied RBPs, waiving the requirement of their binding data. First, we demonstrate the co-evolutionary relationship between RBPs and their RNA partners. We then present a K-nearest neighbors (KNN) based algorithm to infer the sequence preference of an RBP using only the preference information from its homologous RBPs. By benchmarking against several in vitro and in vivo datasets, our proposed method outperforms the existing alternative which uses the closest neighbor's preference on all the datasets. Moreover, it shows comparable performance with two state-of-the-art methods that require the presence of the experimental binding data. Finally, we demonstrate the usage of this method to infer sequence preferences for novel proteins which have no binding preference information available. CONCLUSION: For a poorly studied RBP, the current methods used to determine its binding preference need experimental data, which is expensive and time consuming. Therefore, determining RBP's preference is not practical in many situations. This study provides an economic solution to infer the sequence preference of such protein based on the co-evolution. The source codes and related datasets are available at https://github.com/syang11/KNN . Shu Yang 0009, Junwen Wang, Raymond T. Ng |
BMC Bioinform. | 2 |
| 2016 | Predicting regulatory variants with composite statisticabstractMOTIVATION: Prediction and prioritization of human non-coding regulatory variants is critical for understanding the regulatory mechanisms of disease pathogenesis and promoting personalized medicine. Existing tools utilize functional genomics data and evolutionary information to evaluate the pathogenicity or regulatory functions of non-coding variants. However, different algorithms lead to inconsistent and even conflicting predictions. Combining multiple methods may increase accuracy in regulatory variant prediction. RESULTS: Here, we compiled an integrative resource for predictions from eight different tools on functional annotation of non-coding variants. We further developed a composite strategy to integrate multiple predictions and computed the composite likelihood of a given variant being regulatory variant. Benchmarked by multiple independent causal variants datasets, we demonstrated that our composite model significantly improves the prediction performance. AVAILABILITY AND IMPLEMENTATION: We implemented our model and scoring procedure as a tool, named PRVCS, which is freely available to academic and non-profit usage at http://jjwanglab.org/PRVCS CONTACT: [email protected], [email protected], or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mulin Jun Li, Zipeng Liu, Jiexing Wu, Panwen Wang, Zhengyuan Xia, Pak Chung Sham, Jean-Pierre A. Kocher, Miao-Xin Li, Jun S. Liu, Junwen Wang |
Bioinform. | 13 |
| 2015 | Exploring the function of genetic variants in the non-coding genomic regions: approaches for identifying human regulatory variants affecting gene expressionabstractUnderstanding the genetic basis of human traits/diseases and the underlying mechanisms of how these traits/diseases are affected by genetic variations is critical for public health. Current genome-wide functional genomics data uncovered a large number of functional elements in the noncoding regions of human genome, providing new opportunities to study regulatory variants (RVs). RVs play important roles in transcription factor bindings, chromatin states and epigenetic modifications. Here, we systematically review an array of methods currently used to map RVs as well as the computational approaches in annotating and interpreting their regulatory effects, with emphasis on regulatory single-nucleotide polymorphism. We also briefly introduce experimental methods to validate these functional RVs. Mulin Jun Li, Pak Chung Sham, Junwen Wang |
Briefings Bioinform. | 4 |
| 2014 | CMGRN: a web server for constructing multilevel gene regulatory networks using ChIP-seq and gene expression dataabstractChIP-seq technology provides an accurate characterization of transcription or epigenetic factors binding on genomic sequences. With integration of such ChIP-based and other high-throughput information, it would be dedicated to dissecting cross-interactions among multilevel regulators, genes and biological functions. Here, we devised an integrative web server CMGRN (constructing multilevel gene regulatory networks), to unravel hierarchical interactive networks at different regulatory levels. The newly developed method used the Bayesian network modeling to infer causal interrelationships among transcription factors or epigenetic modifications by using ChIP-seq data. Moreover, it used Bayesian hierarchical model with Gibbs sampling to incorporate binding signals of these regulators and gene expression profile together for reconstructing gene regulatory networks. The example applications indicate that CMGRN provides an effective web-based framework that is able to integrate heterogeneous high-throughput data and to reveal hierarchical 'regulome' and the associated gene expression programs. AVAILABILITY: http://bioinfo.icts.hkbu.edu.hk/cmgrn; http://www.byanbioinfo.org/cmgrn CONTACT: [email protected] or [email protected] Supplementary Information: Supplementary data are available at Bioinformatics online. Daogang Guan, Jiaofang Shao, Youping Deng, Panwen Wang, Zhongying Zhao 0002, Junwen Wang |
Bioinform. | 7 |
| 2014 | FaSD-somatic: a fast and accurate somatic SNV detection algorithm for cancer genome sequencing dataabstractUNLABELLED: Recent advances in high-throughput sequencing technologies have enabled us to sequence large number of cancer samples to reveal novel insights into oncogenetic mechanisms. However, the presence of intratumoral heterogeneity, normal cell contamination and insufficient sequencing depth, together pose a challenge for detecting somatic mutations. Here we propose a fast and an accurate somatic single-nucleotide variations (SNVs) detection program, FaSD-somatic. The performance of FaSD-somatic is extensively assessed on various types of cancer against several state-of-the-art somatic SNV detection programs. Benchmarked by somatic SNVs from either existing databases or de novo higher-depth sequencing data, FaSD-somatic has the best overall performance. Furthermore, FaSD-somatic is efficient, it finishes somatic SNV calling within 14 h on 50X whole genome sequencing data in paired samples. AVAILABILITY AND IMPLEMENTATION: The program, datasets and supplementary files are available at http://jjwanglab.org/FaSD-somatic/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Weixin Wang 0004, Panwen Wang, Ruibang Luo, Maria P. Wong, Tak Wah Lam, Junwen Wang |
Bioinform. | 7 |
| 2014 | DDGni: Dynamic delay gene-network inference from high-temporal data using gapped local alignmentabstractMOTIVATION: Inferring gene-regulatory networks is very crucial in decoding various complex mechanisms in biological systems. Synthesis of a fully functional transcriptional factor/protein from DNA involves series of reactions, leading to a delay in gene regulation. The complexity increases with the dynamic delay induced by other small molecules involved in gene regulation, and noisy cellular environment. The dynamic delay in gene regulation is quite evident in high-temporal live cell lineage-imaging data. Although a number of gene-network-inference methods are proposed, most of them ignore the associated dynamic time delay. RESULTS: Here, we propose DDGni (dynamic delay gene-network inference), a novel gene-network-inference algorithm based on the gapped local alignment of gene-expression profiles. The local alignment can detect short-term gene regulations, that are usually overlooked by traditional correlation and mutual Information based methods. DDGni uses 'gaps' to handle the dynamic delay and non-uniform sampling frequency in high-temporal data, like live cell imaging data. Our algorithm is evaluated on synthetic and yeast cell cycle data, and Caenorhabditis elegans live cell imaging data against other prominent methods. The area under the curve of our method is significantly higher when compared to other methods on all three datasets. AVAILABILITY: The program, datasets and supplementary files are available at http://www.jjwanglab.org/DDGni/. Hari Krishna Yalamanchili, Mulin Jun Li, Jing Qin 0004, Zhongying Zhao 0002, Francis Y. L. Chin, Junwen Wang |
Bioinform. | 7 |
| 2014 | Modeling and Analysis of Care Delivery Services Within Patient Rooms: A System-Theoretic ApproachabstractCare services within the patient rooms are the most critical and time consuming processes in patient care deliveries in emergency department, clinics, and other healthcare facilities. In this paper, we introduce a Markov chain model to study such processes. A closed, parallel, and reentrant network with limited resources is used to model the process. Formulas to evaluate the patient length of stay and staff utilizations are developed. System-theoretic properties are discussed. The extension to non-Markovian scenarios is also investigated. Such a model provides a quantitative tool for healthcare professionals to study and improve patient flow in care deliveries. Junwen Wang, Jingshan Li, Patricia K. Howard |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2013 | Virtual Battery: A Battery Simulation Framework for Electric VehiclesabstractThe battery is one of the most important components in electric vehicles. In this paper, a virtual battery model, which provides a framework of battery simulation for electric vehicles, is introduced. Using such a framework, we can model and simulate the performance of a battery during its usage, such as battery charge, discharge, and idle status, the impacts of internal and external temperature, the manufacturing quality on joints, the cell capacity and balance management, etc. Such a framework can provide a quantitative tool for design and manufacturing engineers to predict the battery performance, investigate the impacts of manufacturing process, and obtain feedback for improvement in battery design, control, and manufacturing processes. Note to Practitioners-Automotive battery manufacturing has become more and more important due to the need of alternative energy source to gasoline powered engines. Although substantial amount of attention has been paid to study both individual battery cells and the battery pack as a whole, a battery model which includes interactions of all its components (cells, joints, external inputs, etc.) is not available, and the impact of manufacturing quality on battery performance has not been investigated. In this paper, a virtual battery simulation framework is developed to evaluate battery performance under different circumstances, involving the issues of cell capacity, temperature, driving profile, the joint (manufacturing) quality, etc. Such a framework can help battery design and manufacturing engineers to evaluate battery performance, investigate the impacts of manufacturing practices, and provide feedback for improvement. Junwen Wang, Jingshan Li, Guoxian Xiao, Stephan R. Biller |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2012 | Reducing Length of Stay in Emergency Department: A Simulation Study at a Community HospitalabstractIn this paper, a simulation model of an emergency department (ED) at a large community hospital, Central Baptist Hospital in Lexington, KY, is developed. Using such a model, we can accurately emulate the patient flow in the ED and carry out sensitivity analysis to determine the most critical process for improvement in quality of care (in terms of patient length of stay). In addition, a what-if analysis is performed to investigate the potential change in operation policies and its impact. Floating nurse, combining registration with triage, mandatory requirement of physician's visit within 30 min, and simultaneous reduction of operation times of some most sensitive procedures can all result in substantial improvement. These recommendations have been submitted to the hospital leadership, and implementations are in progress. Junwen Wang, Jingshan Li, Kathy Tussey, Kay Ross |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 2011 | Correlated evolution of transcription factors and their binding sitesabstractMOTIVATION: The interaction between transcription factor (TF) and transcription factor binding site (TFBS) is essential for gene regulation. Mutation in either the TF or the TFBS may weaken their interaction and thus result in abnormalities. To maintain such vital interaction, a mutation in one of the interacting partners might be compensated by a corresponding mutation in its binding partner during the course of evolution. Confirming this co-evolutionary relationship will guide us in designing protein sequences to target a specific DNA sequence or in predicting TFBS for poorly studied proteins, or even correcting and rescuing disease mutations in clinical applications. RESULTS: Based on six, publicly available, experimentally validated TF-TFBS binding datasets for the basic Helix-Loop-Helix (bHLH) family, Homeo family, High-Mobility Group (HMG) family and Transient Receptor Potential channels (TRP) family, we showed that the evolutions of the TFs and their TFBSs are significantly correlated across eukaryotes. We further developed a mutual information-based method to identify co-evolved protein residues and DNA bases. This research sheds light on the dynamic relationship between TF and TFBS during their evolution. The same principle and strategy can be applied to co-evolutionary studies on protein-DNA interactions in other protein families. AVAILABILITY: All the datasets, scripts and other related files have been made freely available at: http://jjwanglab.org/co-evo. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shu Yang 0009, Hari Krishna Yalamanchili, Kwok-Ming Yao, Pak Chung Sham, Michael Q. Zhang, Junwen Wang |
Bioinform. | 7 |
| 2011 | A passive image authentication scheme for detecting region-duplication forgery with rotation
Guangjie Liu 0001, Junwen Wang, Shiguo Lian |
J. Netw. Comput. Appl. | 2 |
| 2010 | Quality bottleneck transitions in flexible manufacturing systemsabstractIn this paper, we introduce a Markov chain model to evaluate the quality performance in flexible manufacturing systems with batch productions. In such a model, the product quality is a function of the transition probabilities characterizing the changes among good and defective states (where good quality or defective parts are produced during a cycle, respectively). A transition that has the largest impact on quality, i.e., whose improvement will lead to the largest improvement in quality, is defined as the quality bottleneck transition (BN-t). Analytical expressions of sensitivity of quality with respect to transition probabilities are derived. Indicators to identify bottleneck transitions based on the data collected on the factory floor are developed. Numerical experiments show that such indicators have high accuracy in identifying the correct bottlenecks and can be used as an effective tool for quality improvement effort. Finally, a case study at an automotive paint shop to improve quality through quality bottleneck transition identification is introduced. Junwen Wang, Jingshan Li, Jorge Arinez 0001, Stephan R. Biller |
ICRA | 1 |
| 2010 | FastPval: a fast and memory efficient program to calculate very low P-values from empirical distributionabstractMOTIVATION: Resampling methods, such as permutation and bootstrap, have been widely used to generate an empirical distribution for assessing the statistical significance of a measurement. However, to obtain a very low P-value, a large size of resampling is required, where computing speed, memory and storage consumption become bottlenecks, and sometimes become impossible, even on a computer cluster. RESULTS: We have developed a multiple stage P-value calculating program called FastPval that can efficiently calculate very low (up to 10(-9)) P-values from a large number of resampled measurements. With only two input files and a few parameter settings from the users, the program can compute P-values from empirical distribution very efficiently, even on a personal computer. When tested on the order of 10(9) resampled data, our method only uses 52.94% the time used by the conventional method, implemented by standard quicksort and binary search algorithms, and consumes only 0.11% of the memory and storage. Furthermore, our method can be applied to extra large datasets that the conventional method fails to calculate. The accuracy of the method was tested on data generated from Normal, Poison and Gumbel distributions and was found to be no different from the exact ranking approach. AVAILABILITY: The FastPval executable file, the java GUI and source code, and the java web start server with example data and introduction, are available at http://wanglab.hku.hk/pvalue. Mulin Jun Li, Pak Chung Sham, Junwen Wang |
Bioinform. | 3 |
| 2010 | Product Sequencing With Respect to Quality in Flexible Manufacturing Systems With Batch OperationsabstractIn many flexible manufacturing systems, batch production is often adopted to improve product quality. For example, in automotive paint shops, vehicles with same colors are typically grouped into small batches to reduce quality degradation and purge cost due to color change. In this paper, we present an analytical method to evaluate the quality performance of flexible manufacturing systems with batch operations. In addition, we investigate the impact of product sequencing and batch policies on product quality and present some insights to achieve better quality using these policies. Junwen Wang, Jingshan Li, Jorge Arinez 0001, Stephan R. Biller |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2010 | Quality Analysis in Flexible Manufacturing Systems With Batch Productions: Performance Evaluation and Nonmonotonic PropertiesabstractIn this paper, we present an analytical method to evaluate the quality performance of flexible manufacturing systems with batch operations. By using a Markov chain model, a closed formula to quantify the probability of producing a good part is derived and nonmonotonic properties in quality are investigated. Junwen Wang, Jingshan Li, Jorge Arinez 0001, Stephan R. Biller, Ningjian Huang |
IEEE Trans Autom. Sci. Eng. | 1 |
| 2005 | NdPASA: a pairwise sequence alignment server for distantly related proteinsabstractSUMMARY: NdPASA is a web server specifically designed to optimize sequence alignment between distantly related proteins. The program integrates structure information of the template sequence into a global alignment algorithm by employing neighbor-dependent propensities of amino acids as a unique parameter for alignment. NdPASA optimizes alignment by evaluating the likelihood of a residue pair in the query sequence matching against a corresponding residue pair adopting a particular secondary structure in the template sequence. NdPASA is most effective in aligning homologous proteins sharing low percentage of sequence identity. The server is designed to aid homologous protein structure modeling. A PSI-BLAST search engine was implemented to help users identify template candidates that are most appropriate for modeling the query sequences. Junwen Wang, Jin-An Feng |
Bioinform. | 2 |
| 2005 | Generalizations of Markov model to characterize biological sequencesabstractBACKGROUND: The currently used kth order Markov models estimate the probability of generating a single nucleotide conditional upon the immediately preceding (gap = 0) k units. However, this neither takes into account the joint dependency of multiple neighboring nucleotides, nor does it consider the long range dependency with gap > 0. RESULT: We describe a configurable tool to explore generalizations of the standard Markov model. We evaluated whether the sequence classification accuracy can be improved by using an alternative set of model parameters. The evaluation was done on four classes of biological sequences--CpG-poor promoters, all promoters, exons and nucleosome positioning sequences. Using di- and tri-nucleotide as the model unit significantly improved the sequence classification accuracy relative to the standard single nucleotide model. In the case of nucleosome positioning sequences, optimal accuracy was achieved at a gap length of 4. Furthermore in the plot of classification accuracy versus the gap, a periodicity of 10-11 bps was observed which might indicate structural preferences in the nucleosome positioning sequence. The tool is implemented in Java and is available for download at ftp://ftp.pcbi.upenn.edu/GMM/. CONCLUSION: Markov modeling is an important component of many sequence analysis tools. We have extended the standard Markov model to incorporate joint and long range dependencies between the sequence elements. The proposed generalizations of the Markov model are likely to improve the overall accuracy of sequence analysis tools. Junwen Wang, Sridhar Hannenhalli |
BMC Bioinform. | 1 |