Pengyi Yang

dblp:72/5112 · DBLP profile ↗
← Back
31ranked-venue papers
14as first author
8since 2021 · last 2025
0000-0003-1098-3138ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 23 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 5 · 4 first-authorDatabases, data management, data science and information retrieval · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2025 The current landscape and emerging challenges of benchmarking single-cell methods
abstract
With the rapid development of computational methods for single-cell sequencing data, benchmarking serves as a valuable resource. As the number of benchmarking studies surges, it is timely to assess the current state of the field. We conducted a systematic literature search and assessed 282 papers, including all 130 benchmark-only papers from the search and an additional 152 method development papers containing benchmarking. This collective effort provides the most comprehensive quantitative summary of the current landscape of single-cell benchmarking studies. We examine performances across nine broad categories, including often ignored aspects such as role of datasets, robustness of methods and downstream evaluation. Our analysis highlights challenges such as how to effectively combine knowledge across multiple benchmarking studies and in what ways can the community recognize the risk and prevent benchmarking fatigue. This paper highlights the importance of adopting a community-led research paradigm to tackle these challenges and establish best practice standards.
Lijia Yu, Marni Torkel, Yingxin Lin, Pengyi Yang, Terence P. Speed, Shila Ghazanfar, Jean Y. H. Yang
Briefings Bioinform.6
2025 Multi-view gene panel characterization for spatially resolved omics
abstract
Spatially resolved transcriptomics has revolutionized the study of complex tissues by enabling cellular and subcellular resolution. However, targeted spatial technologies depend on pre-selected gene panels, which are typically curated based on prior biological knowledge or specific research hypotheses. While existing methods often focus on optimizing for cell type identification, we argue that effective panel design should also account for transcriptional variation, pathway-level coverage, and minimal gene redundancy. To meet these broader criteria, we developed a two-part framework: (i) panelScope, a gene panel characterization platform that characterizes panels from multiple perspectives, allowing for holistic comparisons of gene panels for custom panel design; and (ii) panelScope-OA, a genetic algorithm that integrates these characterization metrics into a multi-loss function to automate panel optimization. We applied panelScope and panelScope-OA to characterize nine panels across four datasets. Notably, computationally constructed gene panels performed competitively in capturing major cell types when compared to our in-house manually curated panel. However, refined manual curation offered distinct advantages, particularly in capturing minor cell types. Our results demonstrate the utility of panelScope and panelScope-OA by offering quantitative and multi-dimensional insights to support the design of panels tailored to diverse research needs.
Wenze Ding, Akira Nguyen Shaw, Marni Torkel, Cameron J. Turtle, Pengyi Yang, Jean Y. H. Yang
Briefings Bioinform.6
2025 CLUEY enables knowledge-guided clustering and cell type detection from single-cell omics data
abstract
MOTIVATION: Clustering is a fundamental task in single-cell omics data analysis and can significantly impact downstream analyses and biological interpretations. The standard approach involves grouping cells based on their gene expression profiles, followed by annotating each cluster to a cell type using marker genes. However, the number of cell types detected by different clustering methods can vary substantially due to several factors, including the dimension reduction method used and the choice of parameters of the chosen clustering algorithm. These discrepancies can lead to subjective interpretations in downstream analyses, particularly in manual cell type annotation. RESULTS: To address these challenges, we propose CLUEY, a knowledge-guided framework for cell type detection and clustering of single-cell omics data. CLUEY integrates prior biological knowledge into the clustering process, providing guidance on the optimal number of clusters and enhancing the interpretability of results. We apply CLUEY to both unimodal (e.g. scRNA-seq, scATAC-seq) and multimodal datasets (e.g. CITE-seq, SHARE-seq) and demonstrate its effectiveness in providing biologically meaningful clustering outcomes. These results highlight CLUEY on providing the much-needed guidance in clustering analyses of single-cell omics data. AVAILABILITY AND IMPLEMENTATION: CLUEY package is freely available from https://github.com/SydneyBioX/CLUEY.
Carissa Chen, Lijia Yu, Jean Y. H. Yang, Pengyi Yang
Bioinform.5
2024 Interpretable deep learning in single-cell omics
abstract
MOTIVATION: Single-cell omics technologies have enabled the quantification of molecular profiles in individual cells at an unparalleled resolution. Deep learning, a rapidly evolving sub-field of machine learning, has instilled a significant interest in single-cell omics research due to its remarkable success in analysing heterogeneous high-dimensional single-cell omics data. Nevertheless, the inherent multi-layer nonlinear architecture of deep learning models often makes them 'black boxes' as the reasoning behind predictions is often unknown and not transparent to the user. This has stimulated an increasing body of research for addressing the lack of interpretability in deep learning models, especially in single-cell omics data analyses, where the identification and understanding of molecular regulators are crucial for interpreting model predictions and directing downstream experimental validations. RESULTS: In this work, we introduce the basics of single-cell omics technologies and the concept of interpretable deep learning. This is followed by a review of the recent interpretable deep learning models applied to various single-cell omics research. Lastly, we highlight the current limitations and discuss potential future directions.
Manoj M. Wagle, Siqu Long, Carissa Chen, Pengyi Yang
Bioinform.5
2023 Benchmarking of analytical combinations for COVID-19 outcome prediction using single-cell RNA sequencing data
abstract
The advances of single-cell transcriptomic technologies have led to increasing use of single-cell RNA sequencing (scRNA-seq) data in large-scale patient cohort studies. The resulting high-dimensional data can be summarized and incorporated into patient outcome prediction models in several ways; however, there is a pressing need to understand the impact of analytical decisions on such model quality. In this study, we evaluate the impact of analytical choices on model choices, ensemble learning strategies and integrate approaches on patient outcome prediction using five scRNA-seq COVID-19 datasets. First, we examine the difference in performance between using single-view feature space versus multi-view feature space. Next, we survey multiple learning platforms from classical machine learning to modern deep learning methods. Lastly, we compare different integration approaches when combining datasets is necessary. Through benchmarking such analytical combinations, our study highlights the power of ensemble learning, consistency among different learning methods and robustness to dataset normalization when using multiple datasets as the model input.
Shila Ghazanfar, Pengyi Yang, Jean Y. H. Yang
Briefings Bioinform.3
2023 Ensemble deep learning of embeddings for clustering multimodal single-cell omics data
abstract
MOTIVATION: Recent advances in multimodal single-cell omics technologies enable multiple modalities of molecular attributes, such as gene expression, chromatin accessibility, and protein abundance, to be profiled simultaneously at a global level in individual cells. While the increasing availability of multiple data modalities is expected to provide a more accurate clustering and characterization of cells, the development of computational methods that are capable of extracting information embedded across data modalities is still in its infancy. RESULTS: We propose SnapCCESS for clustering cells by integrating data modalities in multimodal single-cell omics data using an unsupervised ensemble deep learning framework. By creating snapshots of embeddings of multimodality using variational autoencoders, SnapCCESS can be coupled with various clustering algorithms for generating consensus clustering of cells. We applied SnapCCESS with several clustering algorithms to various datasets generated from popular multimodal single-cell omics technologies. Our results demonstrate that SnapCCESS is effective and more efficient than conventional ensemble deep learning-based clustering methods and outperforms other state-of-the-art multimodal embedding generation methods in integrating data modalities for clustering cells. The improved clustering of cells from SnapCCESS will pave the way for more accurate characterization of cell identity and types, an essential step for various downstream analyses of multimodal single-cell omics data. AVAILABILITY AND IMPLEMENTATION: SnapCCESS is implemented as a Python package and is freely available from https://github.com/PYangLab/SnapCCESS under the open-source license of GPL-3. The data used in this study are publicly available (see section 'Data availability').
Lijia Yu, Jean Y. H. Yang, Pengyi Yang
Bioinform.4
2022 scFeatures: multi-view representations of single-cell and spatial data for disease outcome prediction
abstract
MOTIVATION: With the recent surge of large-cohort scale single cell research, it is of critical importance that analytical methods can fully utilize the comprehensive characterization of cellular systems that single cell technologies produce to provide insights into samples from individuals. Currently, there is little consensus on the best ways to compress information from the complex data structures of these technologies to summary statistics that represent each sample (e.g. individuals). RESULTS: Here, we present scFeatures, an approach that creates interpretable cellular and molecular representations of single-cell and spatial data at the sample level. We demonstrate that summarizing a broad collection of features at the sample level is both important for understanding underlying disease mechanisms in different experimental studies and for accurately classifying disease status of individuals. AVAILABILITY AND IMPLEMENTATION: scFeatures is publicly available as an R package at https://github.com/SydneyBioX/scFeatures. All data used in this study are publicly available with accession ID reported in the Section 2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yingxin Lin, Ellis Patrick, Pengyi Yang, Jean Y. H. Yang
Bioinform.4
2022 Functional analysis of the stable phosphoproteome reveals cancer vulnerabilities
abstract
MOTIVATION: The advance of mass spectrometry-based technologies enabled the profiling of the phosphoproteomes of a multitude of cell and tissue types. However, current research primarily focused on investigating the phosphorylation dynamics in specific cell types and experimental conditions, whereas the phosphorylation events that are common across cell/tissue types and stable regardless of experimental conditions are, so far, mostly ignored. RESULTS: Here, we developed a statistical framework to identify the stable phosphoproteome across 53 human phosphoproteomics datasets, covering 40 cell/tissue types and 194 conditions/treatments. We demonstrate that the stably phosphorylated sites (SPSs) identified from our statistical framework are evolutionarily conserved, functionally important and enriched in a range of core signaling and gene pathways. Particularly, we show that SPSs are highly enriched in the RNA splicing pathway, an essential cellular process in mammalian cells, and frequently disrupted by cancer mutations, suggesting a link between the dysregulation of RNA splicing and cancer development through mutations on SPSs. AVAILABILITY AND IMPLEMENTATION: The source code for data analysis in this study is available from Github repository https://github.com/PYangLab/SPSs under the open-source license of GPL-3. The data used in this study are publicly available (see Section 2.8). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hani Jieun Kim, Chi N. I. Pang, Pengyi Yang
Bioinform.4
2020 CiteFuse enables multi-modal analysis of CITE-seq data
abstract
MOTIVATION: Multi-modal profiling of single cells represents one of the latest technological advancements in molecular biology. Among various single-cell multi-modal strategies, cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq) allows simultaneous quantification of two distinct species: RNA and cell-surface proteins. Here, we introduce CiteFuse, a streamlined package consisting of a suite of tools for doublet detection, modality integration, clustering, differential RNA and protein expression analysis, antibody-derived tag evaluation, ligand-receptor interaction analysis and interactive web-based visualization of CITE-seq data. RESULTS: We demonstrate the capacity of CiteFuse to integrate the two data modalities and its relative advantage against data generated from single-modality profiling using both simulations and real-world CITE-seq data. Furthermore, we illustrate a novel doublet detection method based on a combined index of cell hashing and transcriptome data. Finally, we demonstrate CiteFuse for predicting ligand-receptor interactions by using multi-modal CITE-seq data. Collectively, we demonstrate the utility and effectiveness of CiteFuse for the integrative analysis of transcriptome and epitope profiles from CITE-seq data. AVAILABILITY AND IMPLEMENTATION: CiteFuse is freely available at http://shiny.maths.usyd.edu.au/CiteFuse/ as an online web service and at https://github.com/SydneyBioX/CiteFuse/ as an R package. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hani Jieun Kim, Yingxin Lin, Thomas Andrew Geddes, Jean Y. H. Yang, Pengyi Yang
Bioinform.5
2019 New Parallel Algorithms for All Pairwise Computation on Large HPC Clusters
abstract
All pairwise computation is defined as performing computation between every pair of the elements in a given dataset. It is often a necessary first step in a number of bioinformatics applications. Many of such applications require multiple terabytes of main memory and take multiple peta floating point operations to complete the computation. Therefore, large HPC clusters are needed to tackle these large-scale computational problems. Conventionally designed parallel algorithms using data partitioning may have a scalability issue, i.e., for a given problem of fixed size the efficiency may decrease if the number of compute nodes is increased (Amdahl's law). In this paper we introduce a new method for parallel algorithm design. Using this method we first design an efficient one-dimensional (1D) ring algorithm and then a two-dimensional (2D) algorithm based on the 1D ring for all pairwise computation. When increasing the compute nodes, instead of reducing the block size, we make multiple copies of the original data blocks in the 1D ring and distribute them across the added compute nodes in the other dimension. By properly organizing the compute nodes the communication overhead can be reduced to a minimum in this two-dimensional setting. Experiments on a Cray XC40 HPC supercomputer show that our new algorithms are very efficient and scalable for large-scale all pairwise computation on large HPC clusters.
Wei Bao 0001, Pengyi Yang, Dong Yuan 0001, Bing Bing Zhou
PDCAT4
2019 Impact of similarity metrics on single-cell RNA-seq data clustering
abstract
Advances in high-throughput sequencing on single-cell gene expressions [single-cell RNA sequencing (scRNA-seq)] have enabled transcriptome profiling on individual cells from complex samples. A common goal in scRNA-seq data analysis is to discover and characterise cell types, typically through clustering methods. The quality of the clustering therefore plays a critical role in biological discovery. While numerous clustering algorithms have been proposed for scRNA-seq data, fundamentally they all rely on a similarity metric for categorising individual cells. Although several studies have compared the performance of various clustering algorithms for scRNA-seq data, currently there is no benchmark of different similarity metrics and their influence on scRNA-seq data clustering. Here, we compared a panel of similarity metrics on clustering a collection of annotated scRNA-seq datasets. Within each dataset, a stratified subsampling procedure was applied and an array of evaluation measures was employed to assess the similarity metrics. This produced a highly reliable and reproducible consensus on their performance assessment. Overall, we found that correlation-based metrics (e.g. Pearson's correlation) outperformed distance-based metrics (e.g. Euclidean distance). To test if the use of correlation-based metrics can benefit the recently published clustering techniques for scRNA-seq data, we modified a state-of-the-art kernel-based clustering algorithm (SIMLR) using Pearson's correlation as a similarity measure and found significant performance improvement over Euclidean distance on scRNA-seq data clustering. These findings demonstrate the importance of similarity metrics in clustering scRNA-seq data and highlight Pearson's correlation as a favourable choice. Further comparison on different scRNA-seq library preparation protocols suggests that they may also affect clustering performance. Finally, the benchmarking framework is available at http://www.maths.usyd.edu.au/u/SMS/bioinformatics/software.html.
Taiyun Kim, Irene Rui Chen, Yingxin Lin, Andy Yi-Yang Wang, Jean Y. H. Yang, Pengyi Yang
Briefings Bioinform.6
2019 scDC: single cell differential composition analysis
abstract
BACKGROUND: Differences in cell-type composition across subjects and conditions often carry biological significance. Recent advancements in single cell sequencing technologies enable cell-types to be identified at the single cell level, and as a result, cell-type composition of tissues can now be studied in exquisite detail. However, a number of challenges remain with cell-type composition analysis - none of the existing methods can identify cell-type perfectly and variability related to cell sampling exists in any single cell experiment. This necessitates the development of method for estimating uncertainty in cell-type composition. RESULTS: We developed a novel single cell differential composition (scDC) analysis method that performs differential cell-type composition analysis via bootstrap resampling. scDC captures the uncertainty associated with cell-type proportions of each subject via bias-corrected and accelerated bootstrap confidence intervals. We assessed the performance of our method using a number of simulated datasets and synthetic datasets curated from publicly available single cell datasets. In simulated datasets, scDC correctly recovered the true cell-type proportions. In synthetic datasets, the cell-type compositions returned by scDC were highly concordant with reference cell-type compositions from the original data. Since the majority of datasets tested in this study have only 2 to 5 subjects per condition, the addition of confidence intervals enabled better comparisons of compositional differences between subjects and across conditions. CONCLUSIONS: scDC is a novel statistical method for performing differential cell-type composition analysis for scRNA-seq data. It uses bootstrap resampling to estimate the standard errors associated with cell-type proportion estimates and performs significance testing through GLM and GLMM models. We have made this method available to the scientific community as part of the scdney package (Single Cell Data Integrative Analysis) R package, available from https://github.com/SydneyBioX/scdney.
Yingxin Lin, John T. Ormerod, Pengyi Yang, Jean Y. H. Yang, Kitty K. Lo
BMC Bioinform.4
2019 Autoencoder-based cluster ensembles for single-cell RNA-seq data analysis
abstract
BACKGROUND: Single-cell RNA-sequencing (scRNA-seq) is a transformative technology, allowing global transcriptomes of individual cells to be profiled with high accuracy. An essential task in scRNA-seq data analysis is the identification of cell types from complex samples or tissues profiled in an experiment. To this end, clustering has become a key computational technique for grouping cells based on their transcriptome profiles, enabling subsequent cell type identification from each cluster of cells. Due to the high feature-dimensionality of the transcriptome (i.e. the large number of measured genes in each cell) and because only a small fraction of genes are cell type-specific and therefore informative for generating cell type-specific clusters, clustering directly on the original feature/gene dimension may lead to uninformative clusters and hinder correct cell type identification. RESULTS: Here, we propose an autoencoder-based cluster ensemble framework in which we first take random subspace projections from the data, then compress each random projection to a low-dimensional space using an autoencoder artificial neural network, and finally apply ensemble clustering across all encoded datasets to generate clusters of cells. We employ four evaluation metrics to benchmark clustering performance and our experiments demonstrate that the proposed autoencoder-based cluster ensemble can lead to substantially improved cell type-specific clusters when applied with both the standard k-means clustering algorithm and a state-of-the-art kernel-based clustering algorithm (SIMLR) designed specifically for scRNA-seq data. Compared to directly using these clustering algorithms on the original datasets, the performance improvement in some cases is up to 100%, depending on the evaluation metric used. CONCLUSIONS: Our results suggest that the proposed framework can facilitate more accurate cell type identification as well as other downstream analyses. The code for creating the proposed autoencoder-based cluster ensemble framework is freely available from https://github.com/gedcom/scCCESS.
Thomas Andrew Geddes, Taiyun Kim, Lihao Nan, James G. Burchfield, Jean Y. H. Yang, Dacheng Tao, Pengyi Yang
BMC Bioinform.7
2019 AdaSampling for Positive-Unlabeled and Label Noise Learning With Bioinformatics Applications
abstract
Class labels are required for supervised learning but may be corrupted or missing in various applications. In binary classification, for example, when only a subset of positive instances is labeled whereas the remaining are unlabeled, positive-unlabeled (PU) learning is required to model from both positive and unlabeled data. Similarly, when class labels are corrupted by mislabeled instances, methods are needed for learning in the presence of class label noise (LN). Here we propose adaptive sampling (AdaSampling), a framework for both PU learning and learning with class LN. By iteratively estimating the class mislabeling probability with an adaptive sampling procedure, the proposed method progressively reduces the risk of selecting mislabeled instances for model training and subsequently constructs highly generalizable models even when a large proportion of mislabeled instances is present in the data. We demonstrate the utilities of proposed methods using simulation and benchmark data, and compare them to alternative approaches that are commonly used for PU learning and/or learning with LN. We then introduce two novel bioinformatics applications where AdaSampling is used to: 1) identify kinase-substrates from mass spectrometry-based phosphoproteomics data and 2) predict transcription factor target genes by integrating various next-generation sequencing data.
Pengyi Yang, John T. Ormerod, Wei Liu 0007, Chendong Ma, Albert Y. Zomaya, Jean Y. H. Yang
IEEE Trans. Cybern.1
2017 Positive unlabeled learning via wrapper-based adaptive sampling
abstract
Learning from positive and unlabeled data frequently occurs in applications where only a subset of positive instances is available while the rest of the data are unlabeled. In such scenarios, often the goal is to create a discriminant model that can accurately classify both positive and negative data by modelling from labeled and unlabeled instances. In this study, we propose an adaptive sampling (AdaSampling) approach that utilises prediction probabilities from a model to iteratively update the training data. Starting with equal prior probabilities for all unlabeled data, our method "wraps" around a predictive model to iteratively update these probabilities to distinguish positive and negative instances in unlabeled data. Subsequently, one or more robust negative set(s) can be drawn from unlabeled data, according to the likelihood of each instance being negative, to train a single classification model or ensemble of models.
Pengyi Yang, Wei Liu 0007, Jean Y. H. Yang
IJCAI1
2017 Integrative analysis identifies co-dependent gene expression regulation of BRG1 and CHD7 at distal regulatory sites in embryonic stem cells
abstract
MOTIVATION: DNA binding proteins such as chromatin remodellers, transcription factors (TFs), histone modifiers and co-factors often bind cooperatively to activate or repress their target genes in a cell type-specific manner. Nonetheless, the precise role of cooperative binding in defining cell-type identity is still largely uncharacterized. RESULTS: Here, we collected and analyzed 214 public datasets representing chromatin immunoprecipitation followed by sequencing (ChIP-Seq) of 104 DNA binding proteins in embryonic stem cell (ESC) lines. We classified their binding sites into those proximal to gene promoters and those in distal regions, and developed a web resource called Proximal And Distal (PAD) clustering to identify their co-localization at these respective regions. Using this extensive dataset, we discovered an extensive co-localization of BRG1 and CHD7 at distal but not proximal regions. The comparison of co-localization sites to those bound by either BRG1 or CHD7 alone showed an enrichment of ESC master TFs binding and active chromatin architecture at co-localization sites. Most notably, our analysis reveals the co-dependency of BRG1 and CHD7 at distal regions on regulating expression of their common target genes in ESC. This work sheds light on cooperative binding of TF binding proteins in regulating gene expression in ESC, and demonstrates the utility of integrative analysis of a manually curated compendium of genome-wide protein binding profiles in our online resource PAD. AVAILABILITY AND IMPLEMENTATION: PAD is freely available at http://pad.victorchang.edu.au/ and its source code is available via an open source GPL 3.0 license at https://github.com/VCCRI/PAD/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengyi Yang, Andrew J. Oldfield, Taiyun Kim, Andrian Yang, Jean Y. H. Yang, Joshua W. K. Ho
Bioinform.1
2016 Positive-unlabeled ensemble learning for kinase substrate prediction from dynamic phosphoproteomics data
abstract
MOTIVATION: Protein phosphorylation is a post-translational modification that underlines various aspects of cellular signaling. A key step to reconstructing signaling networks involves identification of the set of all kinases and their substrates. Experimental characterization of kinase substrates is both expensive and time-consuming. To expedite the discovery of novel substrates, computational approaches based on kinase recognition sequence (motifs) from known substrates, protein structure, interaction and co-localization have been proposed. However, rarely do these methods take into account the dynamic responses of signaling cascades measured from in vivo cellular systems. Given that recent advances in mass spectrometry-based technologies make it possible to quantify phosphorylation on a proteome-wide scale, computational approaches that can integrate static features with dynamic phosphoproteome data would greatly facilitate the prediction of biologically relevant kinase-specific substrates. RESULTS: Here, we propose a positive-unlabeled ensemble learning approach that integrates dynamic phosphoproteomics data with static kinase recognition motifs to predict novel substrates for kinases of interest. We extended a positive-unlabeled learning technique for an ensemble model, which significantly improves prediction sensitivity on novel substrates of kinases while retaining high specificity. We evaluated the performance of the proposed model using simulation studies and subsequently applied it to predict novel substrates of key kinases relevant to insulin signaling. Our analyses show that static sequence motifs and dynamic phosphoproteomics data are complementary and that the proposed integrated model performs better than methods relying only on static information for accurate prediction of kinase-specific substrates. AVAILABILITY AND IMPLEMENTATION: Executable GUI tool, source code and documentation are freely available at https://github.com/PengyiYang/KSP-PUEL. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengyi Yang, Sean J. Humphrey, David E. James, Jean Y. H. Yang, Raja Jothi
Bioinform.1
2016 Highlights from the 11th ISCB Student Council Symposium 2015: Dublin, Ireland. 10 July 2015
abstract
Table of contents A1 Highlights from the eleventh ISCB Student Council Symposium 2015 Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman O1 Prioritizing a drug’s targets using both gene expression and structural similarity Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau O2 Organism specific protein-RNA recognition: A computational analysis of protein-RNA complex structures from different organisms Nagarajan Raju, Sonia Pankaj Chothani, C. Ramakrishnan, Masakazu Sekijima; M. Michael Gromiha O3 Detection of Heterogeneity in Single Particle Tracking Trajectories Paddy J Slator, Nigel J Burroughs O4 3D-NOME: 3D NucleOme Multiscale Engine for data-driven modeling of three-dimensional genome architecture Przemysław Szałaj, Zhonghui Tang, Paul Michalski, Oskar Luo, Xingwang Li, Yijun Ruan, Dariusz Plewczynski O5 A novel feature selection method to extract multiple adjacent solutions for viral genomic sequences classification Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici O6 A Systems Biology Compendium for Leishmania donovani Bart Cuypers, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens O7 Unravelling signal coordination from large scale phosphorylation kinetic data Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic
Katie Wilkins, Mehedi Hassan, Margherita Francescatto, Jakob B. Jespersen, R. Gonzalo Parra, Bart Cuypers, Dan F. DeBlasio, Alexander Junge, Anupama Jigisha, Farzana Rahman, Griet Laenen, Sander Willems, Lieven Thorrez, Yves Moreau, Raju Nagarajan, Sonia P. Chothani, C. Ramakrishnan, Masakazu Sekijima, M. Michael Gromiha, Paddy Slator, Nigel J. Burroughs, Przemyslaw Szalaj, Zhonghui Tang, Paul J. Michalski, Oskar Luo, Xingwang Li 0004, Yijun Ruan, Dariusz Plewczynski, Giulia Fiscon, Emanuel Weitschek, Massimo Ciccozzi, Paola Bertolazzi, Giovanni Felici, Pieter Meysman, Manu Vanaerschot, Maya Berg, Hideo Imamura, Jean-Claude Dujardin, Kris Laukens, Westa Domanova, James R. Krycer, Rima Chaudhuri, Pengyi Yang, Fatemeh Vafaee, Daniel J. Fazakerley, Sean J. Humphrey, David E. James, Zdenka Kuncic
BMC Bioinform.43
2015 Knowledge-Based Analysis for Detecting Key Signaling Events from Time-Series Phosphoproteomics Data
abstract
Cell signaling underlies transcription/epigenetic control of a vast majority of cell-fate decisions. A key goal in cell signaling studies is to identify the set of kinases that underlie key signaling events. In a typical phosphoproteomics study, phosphorylation sites (substrates) of active kinases are quantified proteome-wide. By analyzing the activities of phosphorylation sites over a time-course, the temporal dynamics of signaling cascades can be elucidated. Since many substrates of a given kinase have similar temporal kinetics, clustering phosphorylation sites into distinctive clusters can facilitate identification of their respective kinases. Here we present a knowledge-based CLUster Evaluation (CLUE) approach for identifying the most informative partitioning of a given temporal phosphoproteomics data. Our approach utilizes prior knowledge, annotated kinase-substrate relationships mined from literature and curated databases, to first generate biologically meaningful partitioning of the phosphorylation sites and then determine key kinases associated with each cluster. We demonstrate the utility of the proposed approach on two time-series phosphoproteomics datasets and identify key kinases associated with human embryonic stem cell differentiation and insulin signaling pathway. The proposed approach will be a valuable resource in the identification and characterizing of signaling networks from phosphoproteomics data.
Pengyi Yang, Vivek Jayaswal, Guang Hu 0005, Jean Y. H. Yang, Raja Jothi
PLoS Comput. Biol.1
2014 Direction pathway analysis of large-scale proteomics data reveals novel features of the insulin action pathway
abstract
MOTIVATION: With the advancement of high-throughput techniques, large-scale profiling of biological systems with multiple experimental perturbations is becoming more prevalent. Pathway analysis incorporates prior biological knowledge to analyze genes/proteins in groups in a biological context. However, the hypotheses under investigation are often confined to a 1D space (i.e. up, down, either or mixed regulation). Here, we develop direction pathway analysis (DPA), which can be applied to test hypothesis in a high-dimensional space for identifying pathways that display distinct responses across multiple perturbations. RESULTS: Our DPA approach allows for the identification of pathways that display distinct responses across multiple perturbations. To demonstrate the utility and effectiveness, we evaluated DPA under various simulated scenarios and applied it to study insulin action in adipocytes. A major action of insulin in adipocytes is to regulate the movement of proteins from the interior to the cell surface membrane. Quantitative mass spectrometry-based proteomics was used to study this process on a large-scale. The combined dataset comprises four separate treatments. By applying DPA, we identified that several insulin responsive pathways in the plasma membrane trafficking are only partially dependent on the insulin-regulated kinase Akt. We subsequently validated our findings through targeted analysis of key proteins from these pathways using immunoblotting and live cell microscopy. Our results demonstrate that DPA can be applied to dissect pathway networks testing diverse hypotheses and integrating multiple experimental perturbations. AVAILABILITY AND IMPLEMENTATION: The R package 'directPA' is distributed from CRAN under GNU General Public License (GPL)-3 and can be downloaded from: http://cran.r-project.org/web/packages/directPA/index.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengyi Yang, Ellis Patrick, Shi-Xiong Tan, Daniel J. Fazakerley, James G. Burchfield, Christopher Gribben, Matthew J. Prior, David E. James, Jean Y. H. Yang
Bioinform.1
2014 Sample Subset Optimization Techniques for Imbalanced and Ensemble Learning Problems in Bioinformatics Applications
abstract
Data sampling is a widely used technique in a broad range of machine learning problems. Traditional sampling approaches generally rely on random resampling from a given dataset. However, these approaches do not take into consideration additional information, such as sample quality and usefulness. We recently proposed a data sampling technique, called sample subset optimization (SSO). The SSO technique relies on a cross-validation procedure for identifying and selecting the most useful samples as subsets. In this paper, we describe the application of SSO techniques to imbalanced and ensemble learning problems, respectively. For imbalanced learning, the SSO technique is employed as an under-sampling technique for identifying a subset of highly discriminative samples in the majority class. In ensemble learning, the SSO technique is utilized as a generic ensemble technique where multiple optimized subsets of samples from each class are selected for building an ensemble classifier. We demonstrate the utilities and advantages of the proposed techniques on a variety of bioinformatics applications where class imbalance, small sample size, and noisy data are prevalent.
Pengyi Yang, Paul D. Yoo, Juanita I. Fernando, Bing Bing Zhou, Zili Zhang 0001, Albert Y. Zomaya
IEEE Trans. Cybern.1
2013 Ensemble-Based Wrapper Methods for Feature Selection and Class Imbalance Learning
Pengyi Yang, Wei Liu 0007, Bing Bing Zhou, Sanjay Chawla, Albert Y. Zomaya
PAKDD (1)1
2012 OCAP: an open comprehensive analysis pipeline for iTRAQ
abstract
MOTIVATION: Mass spectrometry-based iTRAQ protein quantification is a high-throughput assay for determining relative protein expressions and identifying disease biomarkers. Processing and analysis of these large and complex data involves a number of distinct components and it is desirable to have a pipeline to efficiently integrate these together. To date, there are limited public available comprehensive analysis pipelines for iTRAQ data and many of these existing pipelines have limited visualization tools and no convenient interfaces with downstream analyses. We have developed a new open source comprehensive iTRAQ analysis pipeline, OCAP, integrating a wavelet-based preprocessing algorithm which provides better peak picking, a new quantification algorithm and a suite of visualizsation tools. OCAP is mainly developed in C++ and is provided as a standalone version (OCAP_standalone) as well as an R package. The R package (OCAP) provides the necessary interfaces with downstream statistical analysis.
Penghao Wang 0002, Pengyi Yang, Jean Y. H. Yang
Bioinform.2
2012 Improving X!Tandem on Peptide Identification from Mass Spectrometry by Self-Boosted Percolator
abstract
A critical component in mass spectrometry (MS)-based proteomics is an accurate protein identification procedure. Database search algorithms commonly generate a list of peptide-spectrum matches (PSMs). The validity of these PSMs is critical for downstream analysis since proteins that are present in the sample are inferred from those PSMs. A variety of postprocessing algorithms have been proposed to validate and filter PSMs. Among them, the most popular ones include a semi-supervised learning (SSL) approach known as Percolator and an empirical modeling approach known as PeptideProphet. However, they are predominantly designed for commercial database search algorithms, i.e., SEQUEST and MASCOT. Therefore, it is highly desirable to extend and optimize those PSM postprocessing algorithms for open source database search algorithms such as X!Tandem. In this paper, we propose a Self-boosted Percolator for postprocessing X!Tandem search results. We find that the SSL algorithm utilized by Percolator depends heavily on the initial ranking of PSMs. Starting with a poor PSM ranking list may cause Percolator to perform suboptimally. By implementing Percolator in a cascade learning manner, we can progressively improve the performance through multiple boost runs, enabling many more PSM identifications without sacrificing false discovery rate (FDR).
Pengyi Yang, Penghao Wang 0002, Bing Bing Zhou, Jean Y. H. Yang
IEEE ACM Trans. Comput. Biol. Bioinform.1
2011 Sample Subset Optimization for Classifying Imbalanced Biological Data
Pengyi Yang, Zili Zhang 0001, Bing Bing Zhou, Albert Y. Zomaya
PAKDD (2)1
2011 Gene-gene interaction filtering with ensemble of filters
abstract
BACKGROUND: Complex diseases are commonly caused by multiple genes and their interactions with each other. Genome-wide association (GWA) studies provide us the opportunity to capture those disease associated genes and gene-gene interactions through panels of SNP markers. However, a proper filtering procedure is critical to reduce the search space prior to the computationally intensive gene-gene interaction identification step. In this study, we show that two commonly used SNP-SNP interaction filtering algorithms, ReliefF and tuned ReliefF (TuRF), are sensitive to the order of the samples in the dataset, giving rise to unstable and suboptimal results. However, we observe that the 'unstable' results from multiple runs of these algorithms can provide valuable information about the dataset. We therefore hypothesize that aggregating results from multiple runs of the algorithm may improve the filtering performance. RESULTS: We propose a simple and effective ensemble approach in which the results from multiple runs of an unstable filter are aggregated based on the general theory of ensemble learning. The ensemble versions of the ReliefF and TuRF algorithms, referred to as ReliefF-E and TuRF-E, are robust to sample order dependency and enable a more informative investigation of data characteristics. Using simulated and real datasets, we demonstrate that both the ensemble of ReliefF and the ensemble of TuRF can generate a much more stable SNP ranking than the original algorithms. Furthermore, the ensemble of TuRF achieved the highest success rate in comparison to many state-of-the-art algorithms as well as traditional χ2-test and odds ratio methods in terms of retaining gene-gene interactions.
Pengyi Yang, Joshua W. K. Ho, Jean Y. H. Yang, Bing Bing Zhou
BMC Bioinform.1
2010 Genetic Algorithm-Based Multi-objective Optimisation for QoS-Aware Web Services Composition
Li Li 0006, Pengyi Yang, Ling Ou, Zili Zhang 0001, Peng Cheng 0011
KSEM2
2010 A dynamic wavelet-based algorithm for pre-processing tandem mass spectrometry data
abstract
MOTIVATION: Mass spectrometry (MS)-based proteomics is one of the most commonly used research techniques for identifying and characterizing proteins in biological and medical research. The identification of a protein is the critical first step in elucidating its biological function. Successful protein identification depends on various interrelated factors, including effective analysis of MS data generated in a proteomic experiment. This analysis comprises several stages, often combined in a pipeline or workflow. The first component of the analysis is known as spectra pre-processing. In this component, the raw data generated by the mass spectrometer is processed to eliminate noise and identify the mass-to-charge ratio (m/z) and intensity for the peaks in the spectrum corresponding to the presence of certain peptides or peptide fragments. Since all downstream analyses depend on the pre-processed data, effective pre-processing is critical to protein identification and characterization. There is a critical need for more robust pre-processing algorithms that perform well on tandem mass spectra under a variety of different conditions and can be easily integrated into sophisticated data analysis pipelines for practical wet-lab applications. RESULT: We have developed a new pre-processing algorithm. Based on wavelet theory, our method uses a dynamic peak model to identify peaks. It is designed to be easily integrated into a complete proteomic analysis workflow. We compared the method with other available algorithms using a reference library of raw MS and tandem MS spectra with known protein composition information. Our pre-processing algorithm results in the identification of significantly more peptides and proteins in the downstream analysis for a given false discovery rate. AVAILABILITY: Software available at: http://www.maths.usyd.edu.au/u/penghao/index.html.
Penghao Wang 0002, Pengyi Yang, Jonathan Arthur, Jean Y. H. Yang
Bioinform.2
2010 A genetic ensemble approach for gene-gene interaction identification
abstract
BACKGROUND: It has now become clear that gene-gene interactions and gene-environment interactions are ubiquitous and fundamental mechanisms for the development of complex diseases. Though a considerable effort has been put into developing statistical models and algorithmic strategies for identifying such interactions, the accurate identification of those genetic interactions has been proven to be very challenging. METHODS: In this paper, we propose a new approach for identifying such gene-gene and gene-environment interactions underlying complex diseases. This is a hybrid algorithm and it combines genetic algorithm (GA) and an ensemble of classifiers (called genetic ensemble). Using this approach, the original problem of SNP interaction identification is converted into a data mining problem of combinatorial feature selection. By collecting various single nucleotide polymorphisms (SNP) subsets as well as environmental factors generated in multiple GA runs, patterns of gene-gene and gene-environment interactions can be extracted using a simple combinatorial ranking method. Also considered in this study is the idea of combining identification results obtained from multiple algorithms. A novel formula based on pairwise double fault is designed to quantify the degree of complementarity. CONCLUSIONS: Our simulation study demonstrates that the proposed genetic ensemble algorithm has comparable identification power to Multifactor Dimensionality Reduction (MDR) and is slightly better than Polymorphism Interaction Analysis (PIA), which are the two most popular methods for gene-gene interaction identification. More importantly, the identification results generated by using our genetic ensemble algorithm are highly complementary to those obtained by PIA and MDR. Experimental results from our simulation studies and real world data application also confirm the effectiveness of the proposed genetic ensemble algorithm, as well as the potential benefits of combining identification results from different algorithms.
Pengyi Yang, Joshua W. K. Ho, Albert Y. Zomaya, Bing Bing Zhou
BMC Bioinform.1
2010 A multi-filter enhanced genetic ensemble system for gene selection and sample classification of microarray data
abstract
BACKGROUND: Feature selection techniques are critical to the analysis of high dimensional datasets. This is especially true in gene selection from microarray data which are commonly with extremely high feature-to-sample ratio. In addition to the essential objectives such as to reduce data noise, to reduce data redundancy, to improve sample classification accuracy, and to improve model generalization property, feature selection also helps biologists to focus on the selected genes to further validate their biological hypotheses. RESULTS: In this paper we describe an improved hybrid system for gene selection. It is based on a recently proposed genetic ensemble (GE) system. To enhance the generalization property of the selected genes or gene subsets and to overcome the overfitting problem of the GE system, we devised a mapping strategy to fuse the goodness information of each gene provided by multiple filtering algorithms. This information is then used for initialization and mutation operation of the genetic ensemble system. CONCLUSION: We used four benchmark microarray datasets (including both binary-class and multi-class classification problems) for concept proving and model evaluation. The experimental results indicate that the proposed multi-filter enhanced genetic ensemble (MF-GE) system is able to improve sample classification accuracy, generate more compact gene subset, and converge to the selection results more quickly. The MF-GE system is very flexible as various combinations of multiple filters and classifiers can be incorporated based on the data characteristics and the user preferences.
Pengyi Yang, Bing Bing Zhou, Zili Zhang 0001, Albert Y. Zomaya
BMC Bioinform.1
2010 A clustering based hybrid system for biomarker selection and sample classification of mass spectrometry data
Pengyi Yang, Zili Zhang 0001, Bing Bing Zhou, Albert Y. Zomaya
Neurocomputing1