Wilson Wen Bin Goh

dblp:66/10932 · DBLP profile ↗
← Back
18ranked-venue papers
3as first author
15since 2021 · last 2025
0000-0003-3863-7501ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Genetic Predictors of Social and Cognitive Outcomes in People with Ultra-High-Risk of Psychosis Using Spiking Neural Networks
abstract
The identification of reliable biomarkers predicting outcomes in ultra-high risk (UHR) psychosis remains a vital challenge in preventive psychiatry. While genetic factors are implicated in psychosis risk, specific markers predicting social and cognitive trajectories have remained elusive. In this 24-month longitudinal study of UHR and Healthy Control (HC) participants, we investigated for the first time the predictive relationship between key immune-related genes, inflammatory response genes, stress response genes, and neurological-related genes with social and cognitive outcomes. Participants underwent comprehensive social and cognitive assessments at baseline and three 6-month intervals. We employed a brain-inspired Spiking Neural Network (SNN) model to map gene-behavior interactions and to identify key genetic predictors that influence both social and cognitive functioning. Additionally, we explore subgroup differences within the UHR population to further understand psychosis risk.
Zohreh Gholami Doborjeh, Balkaran Singh, Alexander Sumich, Maryam Doborjeh, Wilson Wen Bin Goh, Nikola K. Kasabov
IJCNN5
2025 Assessing the impact of batch effect associated missing values on downstream analysis in high-throughput biomedical data
abstract
Batch effect associated missing values (BEAMs) are batch-wide missingness induced from the integration of data with different coverage of biomedical features. BEAMs can present substantial challenges in data analysis. This study investigates how BEAMs impact missing value imputation (MVI) and batch effect (BE) correction algorithms (BECAs). Through simulations and analyses of real-world datasets including the Clinical Proteomic Tumour Analysis Consortium (CPTAC), we evaluated six MVI methods: K-nearest neighbors (KNN), Mean, MinProb, Singular Value Decomposition (SVD), Multivariate Imputation by Chained Equations (MICE), and Random Forest (RF), with ComBat and limma as the BECAs. We demonstrated that BEAMs strongly affect MVI performance, resulting in inaccurate imputed values, inflated significant P-values, and compromised BE correction. KNN, SVD, and RF were particularly prone to propagating random signals, resulting in false statistical confidence. While imputation with Mean and MinProb were less detrimental, artifacts were nonetheless introduced. Furthermore, the detrimental effect of BEAMs increased in parallel with its severity in the data. Our findings highlight the necessity of comprehensive assessments and tailored strategies to handle BEAMs in multi-batch datasets to ensure reliable data analysis and interpretation. Future work should investigate more advanced simulations and a variety of dedicated MVI methods to robustly address BEAMs.
Harvard Wai Hann Hui, Wei Xin Chan, Wilson Wen Bin Goh
Briefings Bioinform.3
2025 Protrec2: tissue-specific network-based missing protein recovery method
abstract
Despite technological advances, missing proteins remain a challenge in proteomics, obscuring proteins that are biologically or clinically important. We present Protrec2, a probabilistic framework that integrates tissue-specific protein complex annotations with Bayesian inference to recover unreported but biologically present proteins. We benchmarked Protrec2 on HeLa and A549-derived proteomes under "upper-bound" and "lower-bound" scenarios, reflecting distinct but complementary real-world use cases. In upper-bound evaluations, Protrec2 consistently outperformed state-of-the-art methods such as PROTein RECovery, Functional Class Scoring, Hypergeometric Enrichment, and Gene Set Enrichment Analysis, achieving the highest recovery rates: up to 98.4% in A549 and 96.5% in HeLa and validating 650 and 453 proteins, respectively. In lower-bound evaluations, Protrec2 maintained superior precision, validating over 90% of its predicted proteins in the A549 dataset and 74.6% in HeLa, while other methods exhibited significant performance drops. We applied Protrec2 to six matched lung tumor-normal pairs and validated predictions against CPTAC. Over 85% of predicted proteins were supported, with cancer-specific proteins mostly upregulated and normal-exclusive ones downregulated. Frequently recovered proteins (e.g. P4HA3, SNX1, HIP1R, NOS2) are known to play key roles in lung cancer, highlighting the biological and clinical relevance of Protrec2. These findings establish Protrec2 as a robust, biologically grounded tool for missing protein recovery, with broad applicability in discovery proteomics and translational research.
Weijia Kong, Wilson Wen Bin Goh, Limsoon Wong
Briefings Bioinform.2
2025 Exploiting the similarity of dissimilarities for biomedical applications and enhanced machine learning
abstract
The "similarity of dissimilarities" is an emerging paradigm in biomedical science with significant implications for protein function prediction, machine learning (ML), and personalized medicine. In protein function prediction, recognizing dissimilarities alongside similarities provides a more detailed understanding of evolutionary processes, allowing for a deeper exploration of regions that influence biological functionality. For ML models, incorporating dissimilarity measures helps avoid misleading results caused by highly correlated or similar data, addressing confounding issues like the Doppelgänger Effect. This leads to more accurate insights and a stronger understanding of complex biological systems. In the realm of personalized AI and precision medicine, the importance of dissimilarities is paramount. Personalized AI builds local models for each sample by identifying a network of neighboring samples. However, if the neighboring samples are too similar, it becomes difficult to identify factors critical to disease onset for the individual, limiting the effectiveness of personalized interventions or treatments. This paper discusses the "similarity of dissimilarities" concept, using protein function prediction, ML, and personalized AI as key examples. Integrating this approach into an analysis allows for the design of better, more meaningful experiments and the development of smarter validation methods, ensuring that the models learn in a meaningful way.
Mohammad Neamul Kabir, Li Rong Wang, Wilson Wen Bin Goh
PLoS Comput. Biol.3
2024 Thinking points for effective batch correction on biomedical data
abstract
Batch effects introduce significant variability into high-dimensional data, complicating accurate analysis and leading to potentially misleading conclusions if not adequately addressed. Despite technological and algorithmic advancements in biomedical research, effectively managing batch effects remains a complex challenge requiring comprehensive considerations. This paper underscores the necessity of a flexible and holistic approach for selecting batch effect correction algorithms (BECAs), advocating for proper BECA evaluations and consideration of artificial intelligence-based strategies. We also discuss key challenges in batch effect correction, including the importance of uncovering hidden batch factors and understanding the impact of design imbalance, missing values, and aggressive correction. Our aim is to provide researchers with a robust framework for effective batch effects management and enhancing the reliability of high-dimensional data analyses.
Harvard Wai Hann Hui, Weijia Kong, Wilson Wen Bin Goh
Briefings Bioinform.3
2024 OLB-AC: toward optimizing ligand bioactivities through deep graph learning and activity cliffs
abstract
MOTIVATION: Deep graph learning (DGL) has been widely employed in the realm of ligand-based virtual screening. Within this field, a key hurdle is the existence of activity cliffs (ACs), where minor chemical alterations can lead to significant changes in bioactivity. In response, several DGL models have been developed to enhance ligand bioactivity prediction in the presence of ACs. Yet, there remains a largely unexplored opportunity within ACs for optimizing ligand bioactivity, making it an area ripe for further investigation. RESULTS: We present a novel approach to simultaneously predict and optimize ligand bioactivities through DGL and ACs (OLB-AC). OLB-AC possesses the capability to optimize ligand molecules located near ACs, providing a direct reference for optimizing ligand bioactivities with the matching of original ligands. To accomplish this, a novel attentive graph reconstruction neural network and ligand optimization scheme are proposed. Attentive graph reconstruction neural network reconstructs original ligands and optimizes them through adversarial representations derived from their bioactivity prediction process. Experimental results on nine drug targets reveal that out of the 667 molecules generated through OLB-AC optimization on datasets comprising 974 low-activity, noninhibitor, or highly toxic ligands, 49 are recognized as known highly active, inhibitor, or nontoxic ligands beyond the datasets' scope. The 27 out of 49 matched molecular pairs generated by OLB-AC reveal novel transformations not present in their training sets. The adversarial representations employed for ligand optimization originate from the gradients of bioactivity predictions. Therefore, we also assess OLB-AC's prediction accuracy across 33 different bioactivity datasets. Results show that OLB-AC achieves the best Pearson correlation coefficient (r2) on 27/33 datasets, with an average improvement of 7.2%-22.9% against the state-of-the-art bioactivity prediction methods. AVAILABILITY AND IMPLEMENTATION: The code and dataset developed in this work are available at github.com/Yueming-Yin/OLB-AC.
Yueming Yin, Haifeng Hu 0004, Jitao Yang, Chun Ye, Wilson Wen Bin Goh, Adams Wai-Kin Kong
Bioinform.5
2024 Ten quick tips for ensuring machine learning model validity
abstract
Artificial Intelligence (AI) and Machine Learning (ML) models are increasingly deployed on biomedical and health data to shed insights on biological mechanism, predict disease outcomes, and support clinical decision-making.However, ensuring model validity is challenging.The 10 quick tips described here discuss useful practices on how to check AI/ ML models from 2 perspectives-the user and the developer. IntroductionAU : Pleaseconfirmthatallheadinglevelsarerepresentedcorrectly:The rapid advancement of Machine Learning (ML) and Artificial Intelligence (AI) technologies has sparked a transformative revolution across diverse domains.The convergence of sophisticated algorithms, powerful computing capabilities, and an abundance of data has propelled these technologies to the forefront of innovation, significantly impacting fields such as biomedicine, health, and technology.The increasing importance of ML and AI can be attributed to their unparalleled ability to decipher complex patterns [1], extract valuable insights [2], and automate decision-making processes [3].AI/ML models are increasingly deployed on biomedical and health data.These models can be used to shed insights on biological mechanisms, predict disease outcomes, and support clinical decision-making.We see some notable successes in, for example, protein structure prediction [4] and in clinical decision support [5], but there have also been challenges and less stellar outcomes.For example, in drug target prediction, IBM Watson did not live up to expectations in streamlining and accelerating the drug discovery process.And in meta-analysis, AI models were not able to yield quality explanations due to issues such as random feature substitutability [6][7][8] and the existence of many high-performing models in the Rashomon set [9][10][11].AI and ML are, ultimately, tools.The effectiveness of these tools depends on how well the human user is capable of building and exploiting them [12].Current literature provides general guidelines on using ML models in different areas like chemical science, COVID-19 data, etc. [13-16], which focuses more on input data, leakage, reproducibility, class imbalance,
Wilson Wen Bin Goh, Mohammad Neamul Kabir, Sehwan Yoo, Limsoon Wong
PLoS Comput. Biol.1
2023 Mosaic LSM: A Liquid State Machine Approach for Multimodal Longitudinal Data Analysis
abstract
In this paper, we present a novel Liquid State Machine (LSM) based approach for modelling of multimodal longitudinal data: the Mosaic LSM. Our model harnesses the strengths of multiple LSMs, each designed to capture the temporal patterns of a specific data modality. This temporal information is then added to the raw data to create a composite representation that encompasses both the multimodal and the longitudinal aspects of the data. We demonstrate the performance of our approach on a real-world dataset that contains clinical, cognitive, and genetic modalities with the aim of predicting the Ultra-High Risk (UHR) status in individuals, six months in advance. Our results show that the Mosaic LSM outperforms traditional machine learning models, achieving an outstanding Matthew's Correlation Coefficient of 0.84 and prediction accuracy of 92.4%. Overall, our work highlights the potential of Mosaic LSM as a powerful tool for disease prognosis, and its ability to leverage both the multimodality and temporality of the data to improve performance.
Sugam Budhraja, Balkaran Singh, Maryam Doborjeh, Zohreh Gholami Doborjeh, Samuel Tan, Edmund M.-K. Lai, Wilson Wen Bin Goh, Nikola K. Kasabov
IJCNN7
2023 Filter and Wrapper Stacking Ensemble (FWSE): a robust approach for reliable biomarker discovery in high-dimensional omics data
abstract
Selecting informative features, such as accurate biomarkers for disease diagnosis, prognosis and response to treatment, is an essential task in the field of bioinformatics. Medical data often contain thousands of features and identifying potential biomarkers is challenging due to small number of samples in the data, method dependence and non-reproducibility. This paper proposes a novel ensemble feature selection method, named Filter and Wrapper Stacking Ensemble (FWSE), to identify reproducible biomarkers from high-dimensional omics data. In FWSE, filter feature selection methods are run on numerous subsets of the data to eliminate irrelevant features, and then wrapper feature selection methods are applied to rank the top features. The method was validated on four high-dimensional medical datasets related to mental illnesses and cancer. The results indicate that the features selected by FWSE are stable and statistically more significant than the ones obtained by existing methods while also demonstrating biological relevance. Furthermore, FWSE is a generic method, applicable to various high-dimensional datasets in the fields of machine intelligence and bioinformatics.
Sugam Budhraja, Maryam Doborjeh, Balkaran Singh, Samuel Tan, Zohreh Gholami Doborjeh, Edmund M.-K. Lai, Alexander Merkin, Jimmy Lee, Wilson Wen Bin Goh, Nikola K. Kasabov
Briefings Bioinform.9
2023 ProJect: a powerful mixed-model missing value imputation method
abstract
Missing values (MVs) can adversely impact data analysis and machine-learning model development. We propose a novel mixed-model method for missing value imputation (MVI). This method, ProJect (short for Protein inJection), is a powerful and meaningful improvement over existing MVI methods such as Bayesian principal component analysis (PCA), probabilistic PCA, local least squares and quantile regression imputation of left-censored data. We rigorously tested ProJect on various high-throughput data types, including genomics and mass spectrometry (MS)-based proteomics. Specifically, we utilized renal cancer (RC) data acquired using DIA-SWATH, ovarian cancer (OC) data acquired using DIA-MS, bladder (BladderBatch) and glioblastoma (GBM) microarray gene expression dataset. Our results demonstrate that ProJect consistently performs better than other referenced MVI methods. It achieves the lowest normalized root mean square error (on average, scoring 45.92% less error in RC_C, 27.37% in RC_full, 29.22% in OC, 23.65% in BladderBatch and 20.20% in GBM relative to the closest competing method) and the Procrustes sum of squared error (Procrustes SS) (exhibits 79.71% less error in RC_C, 38.36% in RC full, 18.13% in OC, 74.74% in BladderBatch and 30.79% in GBM compared to the next best method). ProJect also leads with the highest correlation coefficient among all types of MV combinations (0.64% higher in RC_C, 0.24% in RC full, 0.55% in OC, 0.39% in BladderBatch and 0.27% in GBM versus the second-best performing method). ProJect's key strength is its ability to handle different types of MVs commonly found in real-world data. Unlike most MVI methods that are designed to handle only one type of MV, ProJect employs a decision-making algorithm that first determines if an MV is missing at random or missing not at random. It then employs targeted imputation strategies for each MV type, resulting in more accurate and reliable imputation outcomes. An R implementation of ProJect is available at https://github.com/miaomiao6606/ProJect.
Weijia Kong, Bertrand Jern Han Wong, Harvard Wai Hann Hui, Kai Peng Lim, Limsoon Wong, Wilson Wen Bin Goh
Briefings Bioinform.7
2023 PROSE: phenotype-specific network signatures from individual proteomic samples
abstract
Proteomic studies characterize the protein composition of complex biological samples. Despite recent advancements in mass spectrometry instrumentation and computational tools, low proteome coverage and interpretability remains a challenge. To address this, we developed Proteome Support Vector Enrichment (PROSE), a fast, scalable and lightweight pipeline for scoring proteins based on orthogonal gene co-expression network matrices. PROSE utilizes simple protein lists as input, generating a standard enrichment score for all proteins, including undetected ones. In our benchmark with 7 other candidate prioritization techniques, PROSE shows high accuracy in missing protein prediction, with scores correlating strongly to corresponding gene expression data. As a further proof-of-concept, we applied PROSE to a reanalysis of the Cancer Cell Line Encyclopedia proteomics dataset, where it captures key phenotypic features, including gene dependency. We lastly demonstrated its applicability on a breast cancer clinical dataset, showing clustering by annotated molecular subtype and identification of putative drivers of triple-negative breast cancer. PROSE is available as a user-friendly Python module from https://github.com/bwbio/PROSE.
Bertrand Jern Han Wong, Weijia Kong, Wilson Wen Bin Goh
Briefings Bioinform.4
2023 ProInfer: An interpretable protein inference tool leveraging on biological networks
abstract
In mass spectrometry (MS)-based proteomics, protein inference from identified peptides (protein fragments) is a critical step. We present ProInfer (Protein Inference), a novel protein assembly method that takes advantage of information in biological networks. ProInfer assists recovery of proteins supported only by ambiguous peptides (a peptide which maps to more than one candidate protein) and enhances the statistical confidence for proteins supported by both unique and ambiguous peptides. Consequently, ProInfer rescues weakly supported proteins thereby improving proteome coverage. Evaluated across THP1 cell line, lung cancer and RAW267.4 datasets, ProInfer always infers the most numbers of true positives, in comparison to mainstream protein inference tools Fido, EPIFANY and PIA. ProInfer is also adept at retrieving differentially expressed proteins, signifying its usefulness for functional analysis and phenotype profiling. Source codes of ProInfer are available at https://github.com/PennHui2016/ProInfer.
Limsoon Wong, Wilson Wen Bin Goh
PLoS Comput. Biol.3
2023 Transfer Learning of Fuzzy Spatio-Temporal Rules in a Brain-Inspired Spiking Neural Network Architecture: A Case Study on Spatio-Temporal Brain Data
abstract
The article demonstrates for the first time that a brain-inspired spiking neural network (SNN) architecture can be used not only to learn spatio-temporal data, but also to extract fuzzy spatio-temporal rules from such data and to update these rules incrementally in a transfer learning mode. We propose a method, where a SNN model learns incrementally new time-space data related to new classes/tasks/categories, always utilizing some previously learned knowledge, and presents the evolved knowledge as fuzzy spatio-temporal rules. Similarly, to how the brain manifests transfer learning, these SNN models do not need to be restricted in number of layers and neurons in each layer as they adopt self-organizing learning principles. The continuously evolved fuzzy rules from spatio-temporal data are interpretable for a better understanding of the processes that generate the data. The proposed method is based on a brain-inspired SNN architecture NeuCube, which is structured according to a brain three-dimensional structural template. It is illustrated on tasks of incremental and transfer learning and knowledge transfer using spatio-temporal data measuring brain activity, when subjects are performing tasks in space and time. The method is a general one and opens the field to create new types of adaptable and explainable spatio-temporal learning systems across domain areas.
Nikola K. Kasabov, Yongyao Tan, Maryam Doborjeh, Enmei Tu, Jie Yang 0002, Wilson Wen Bin Goh, Jimmy Lee
IEEE Trans. Fuzzy Syst.6
2022 A novel pipeline for computerized mouse spermatogenesis staging
abstract
MOTIVATION: Differentiating 12 stages of the mouse seminiferous epithelial cycle is vital towards understanding the dynamic spermatogenesis process. However, it is challenging since two adjacent spermatogenic stages are morphologically similar. Distinguishing Stages I-III from Stages IV-V is important for histologists to understand sperm development in wildtype mice and spermatogenic defects in infertile mice. To achieve this, we propose a novel pipeline for computerized spermatogenesis staging (CSS). RESULTS: The CSS pipeline comprises four parts: (i) A seminiferous tubule segmentation model is developed to extract every single tubule; (ii) A multi-scale learning (MSL) model is developed to integrate local and global information of a seminiferous tubule to distinguish Stages I-V from Stages VI-XII; (iii) a multi-task learning (MTL) model is developed to segment the multiple testicular cells for Stages I-V without an exhaustive requirement for manual annotation; (iv) A set of 204D image-derived features is developed to discriminate Stages I-III from Stages IV-V by capturing cell-level and image-level representation. Experimental results suggest that the proposed MSL and MTL models outperform classic single-scale and single-task models when manual annotation is limited. In addition, the proposed image-derived features are discriminative between Stages I-III and Stages IV-V. In conclusion, the CSS pipeline can not only provide histologists with a solution to facilitate quantitative analysis for spermatogenesis stage identification but also help them to uncover novel computerized image-derived biomarkers. AVAILABILITY AND IMPLEMENTATION: https://github.com/jydada/CSS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Haoda Lu, Min Zang, Gabriel Pik Liang Marini, Xiangxue Wang, Yiping Jiao, Nianfei Ao, Ong Kok Haur, Xinmi Huo, Longjie Li 0004, Eugene Yujun Xu, Wilson Wen Bin Goh, Weimiao Yu, Jun Xu 0005
Bioinform.11
2021 Predicting Student Performance in Experiential Education
Lejia Lin, Leonard Wee Liat Tan, Nicole Hui Lin Kan, Ooi Kiang Tan, Chun Chau Sze, Wilson Wen Bin Goh
DEXA (1)6
2019 Advanced bioinformatics methods for practical applications in proteomics
abstract
Mass spectrometry (MS)-based proteomics has undergone rapid advancements in recent years, creating challenging problems for bioinformatics. We focus on four aspects where bioinformatics plays a crucial role (and proteomics is needed for clinical application): peptide-spectra matching (PSM) based on the new data-independent acquisition (DIA) paradigm, resolving missing proteins (MPs), dealing with biological and technical heterogeneity in data and statistical feature selection (SFS). DIA is a brute-force strategy that provides greater width and depth but, because it indiscriminately captures spectra such that signal from multiple peptides is mixed, getting good PSMs is difficult. We consider two strategies: simplification of DIA spectra to pseudo-data-dependent acquisition spectra or, alternatively, brute-force search of each DIA spectra against known reference libraries. The MP problem arises when proteins are never (or inconsistently) detected by MS. When observed in at least one sample, imputation methods can be used to guess the approximate protein expression level. If never observed at all, network/protein complex-based contextualization provides an independent prediction platform. Data heterogeneity is a difficult problem with two dimensions: technical (batch effects), which should be removed, and biological (including demography and disease subpopulations), which should be retained. Simple normalization is seldom sufficient, while batch effect-correction algorithms may create errors. Batch effect-resistant normalization methods are a viable alternative. Finally, SFS is vital for practical applications. While many methods exist, there is no best method, and both upstream (e.g. normalization) and downstream processing (e.g. multiple-testing correction) are performance confounders. We also discuss signal detection when class effects are weak.
Wilson Wen Bin Goh, Limsoon Wong
Briefings Bioinform.1
2013 Random forests on Hadoop for genome-wide association studies of multivariate neuroimaging phenotypes
abstract
MOTIVATION: Multivariate quantitative traits arise naturally in recent neuroimaging genetics studies, in which both structural and functional variability of the human brain is measured non-invasively through techniques such as magnetic resonance imaging (MRI). There is growing interest in detecting genetic variants associated with such multivariate traits, especially in genome-wide studies. Random forests (RFs) classifiers, which are ensembles of decision trees, are amongst the best performing machine learning algorithms and have been successfully employed for the prioritisation of genetic variants in case-control studies. RFs can also be applied to produce gene rankings in association studies with multivariate quantitative traits, and to estimate genetic similarities measures that are predictive of the trait. However, in studies involving hundreds of thousands of SNPs and high-dimensional traits, a very large ensemble of trees must be inferred from the data in order to obtain reliable rankings, which makes the application of these algorithms computationally prohibitive. RESULTS: We have developed a parallel version of the RF algorithm for regression and genetic similarity learning tasks in large-scale population genetic association studies involving multivariate traits, called PaRFR (Parallel Random Forest Regression). Our implementation takes advantage of the MapReduce programming model and is deployed on Hadoop, an open-source software framework that supports data-intensive distributed applications. Notable speed-ups are obtained by introducing a distance-based criterion for node splitting in the tree estimation process. PaRFR has been applied to a genome-wide association study on Alzheimer's disease (AD) in which the quantitative trait consists of a high-dimensional neuroimaging phenotype describing longitudinal changes in the human brain structure. PaRFR provides a ranking of SNPs associated to this trait, and produces pair-wise measures of genetic proximity that can be directly compared to pair-wise measures of phenotypic proximity. Several known AD-related variants have been identified, including APOE4 and TOMM40. We also present experimental evidence supporting the hypothesis of a linear relationship between the number of top-ranked mutated states, or frequent mutation patterns, and an indicator of disease severity. AVAILABILITY: The Java codes are freely available at http://www2.imperial.ac.uk/~gmontana.
Yue Wang 0006, Wilson Wen Bin Goh, Limsoon Wong, Giovanni Montana
BMC Bioinform.2
2012 The role of miRNAs in complex formation and control
abstract
UNLABELLED: microRibonucleic acid (miRNAs) are small regulatory molecules that act by mRNA degradation or via translational repression. Although many miRNAs are ubiquitously expressed, a small subset have differential expression patterns that may give rise to tissue-specific complexes. MOTIVATION: This work studies gene targeting patterns amongst miRNAs with differential expression profiles, and links this to control and regulation of protein complexes. RESULTS: We find that, when a pair of miRNAs are not expressed in the same tissues, there is a higher tendency for them to target the direct partners of the same hub proteins. At the same time, they also avoid targeting the same set of hub-spokes. Moreover, the complexes corresponding to these hub-spokes tend to be specific and nonoverlapping. This suggests that the effect of miRNAs on the formation of complexes is specific.
Wilson Wen Bin Goh, Hirotaka Oikawa, Judy Chia Ghee Sng, Marek J. Sergot, Limsoon Wong
Bioinform.1