Jean Y. H. Yang

dblp:27/5876 · also Jean Yee Hwa Yang, Yee Hwa Yang · DBLP profile ↗
← Back
41ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0002-5271-2603ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 37 · 2 first-author · 12 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2
YearPublicationVenuePosition
2026 Estimating tumour immune infiltration: methodological convergence across histology and spatial technologies
abstract
Estimating tumour immune infiltration is critical for understanding cancer biology and predicting patient response to surgery and immunotherapy. A wide array of experimental platforms supports various computational approaches for quantifying tumour-infiltrating lymphocytes and immune infiltration level within the tumour microenvironment, including traditional histopathology, immunohistochemistry, AI-based digital pathology, bulk RNA sequencing, and spatial omics now. Although numerous technologies are available to quantify immune infiltration, important questions remain about which approaches are most suitable and how best to guide platform and methodological choices. In this review, we provide a comprehensive overview of how each platform estimates tumour immune profile coupled with their corresponding computational approaches, followed by a comparative discussion on technological resolution, cell-type specificity, spatial context, and clinical interpretability. We also discuss emerging trends in multimodal data integration, including mapping-based and fusion-based strategies. Together, our review underscores both the methodological opportunities and the translational potential of diverse immune infiltration estimation strategies, guiding the design of more actionable immune profiling strategies.
Beilei Bian, Jean Y. H. Yang
Briefings Bioinform.3
2026 Improving lesion segmentation in medical images by global and regional feature compensation
abstract
• Dual-feature compensation framework to improve medical image segmentation of challenging lesions. • Addresses the loss of detailed global features caused by downsampling. • Uses SSL residual maps to provide additional pixel-level feature representations. • Applies patch-based cross-attention to integrate SSL residual maps, targeting likely lesion regions. Automated lesion segmentation of medical images has made tremendous improvements in recent years due to deep learning advancements. However, accurately capturing fine-grained global and regional feature representations remains a challenge. Many existing methods achieve suboptimal performance in complex lesion segmentation due to information loss during typical downsampling operations and insufficient capture of either regional or global features. To address these issues, we propose the Global and Regional Compensation Segmentation Framework (GRCSF), which introduces two key innovations: the Global Compensation Unit (GCU) and the Region Compensation Unit (RCU). The proposed GCU addresses resolution loss in the U-shaped backbone by preserving global contextual features and fine-grained details during multiscale downsampling. Meanwhile, the RCU introduces a self-supervised learning (SSL) residual map generated by Masked Autoencoders (MAE), obtained as pixel-wise differences between reconstructed and original images, to highlight regions with potential lesions. These SSL residual maps guide precise lesion localization and segmentation through a patch-based cross-attention mechanism that integrates regional spatial and pixel-level features. Additionally, the RCU incorporates patch-level importance scoring to enhance feature fusion by leveraging global spatial information from the backbone. Experiments on three publicly available medical image segmentation datasets, including brain stroke lesion, lung tumor and coronary artery calcification datasets, demonstrate that our GRCSF outperforms state-of-the-art methods, confirming its effectiveness across diverse lesion types and its potential as a generalizable lesion segmentation solution.
Jean Y. H. Yang, Jinman Kim
Pattern Recognit.3
2025 The current landscape and emerging challenges of benchmarking single-cell methods
abstract
With the rapid development of computational methods for single-cell sequencing data, benchmarking serves as a valuable resource. As the number of benchmarking studies surges, it is timely to assess the current state of the field. We conducted a systematic literature search and assessed 282 papers, including all 130 benchmark-only papers from the search and an additional 152 method development papers containing benchmarking. This collective effort provides the most comprehensive quantitative summary of the current landscape of single-cell benchmarking studies. We examine performances across nine broad categories, including often ignored aspects such as role of datasets, robustness of methods and downstream evaluation. Our analysis highlights challenges such as how to effectively combine knowledge across multiple benchmarking studies and in what ways can the community recognize the risk and prevent benchmarking fatigue. This paper highlights the importance of adopting a community-led research paradigm to tackle these challenges and establish best practice standards.
Lijia Yu, Marni Torkel, Yingxin Lin, Pengyi Yang, Terence P. Speed, Shila Ghazanfar, Jean Y. H. Yang
Briefings Bioinform.9
2025 Multi-view gene panel characterization for spatially resolved omics
abstract
Spatially resolved transcriptomics has revolutionized the study of complex tissues by enabling cellular and subcellular resolution. However, targeted spatial technologies depend on pre-selected gene panels, which are typically curated based on prior biological knowledge or specific research hypotheses. While existing methods often focus on optimizing for cell type identification, we argue that effective panel design should also account for transcriptional variation, pathway-level coverage, and minimal gene redundancy. To meet these broader criteria, we developed a two-part framework: (i) panelScope, a gene panel characterization platform that characterizes panels from multiple perspectives, allowing for holistic comparisons of gene panels for custom panel design; and (ii) panelScope-OA, a genetic algorithm that integrates these characterization metrics into a multi-loss function to automate panel optimization. We applied panelScope and panelScope-OA to characterize nine panels across four datasets. Notably, computationally constructed gene panels performed competitively in capturing major cell types when compared to our in-house manually curated panel. However, refined manual curation offered distinct advantages, particularly in capturing minor cell types. Our results demonstrate the utility of panelScope and panelScope-OA by offering quantitative and multi-dimensional insights to support the design of panels tailored to diverse research needs.
Wenze Ding, Akira Nguyen Shaw, Marni Torkel, Cameron J. Turtle, Pengyi Yang, Jean Y. H. Yang
Briefings Bioinform.7
2025 eNODAL: an experimentally guided nutriomics data clustering method to unravel complex drug-diet interactions
abstract
Unraveling the complex interplay between nutrients and drugs via their effects on "omics" features could revolutionize our fundamental understanding of nutritional physiology, personalized nutrition, and, ultimately, human health span. Experimental studies in nutrition are starting to use large-scale "omics" experiments to pick apart the effects of such interacting factors. However, the high dimensionality of the omics features, coupled with complex fully factorial experimental designs, poses a challenge to the analysis. Current strategies for analyzing such types of data are based on between-feature correlations. However, these techniques risk overlooking important signals that arise from the experimental design and produce clusters that are hard to interpret. We present a novel approach for analyzing high-dimensional outcomes in nutriomics experiments, termed experiment-guided NutriOmics DatA cLustering ('eNODAL'). This three-step hybrid framework takes advantage of both Analysis of Variance (ANOVA)-type analyses and unsupervised learning methods to extract maximum information from experimental nutriomics studies. First, eNODAL categorizes the omics features into interpretable groups based on the significance of response to the different experimental variables using an ANOVA-like test. Such groups may include the main effects of a nutritional intervention and drug exposure or their interaction. Second, consensus clustering is performed within each interpretable group to further identify subclusters of features with similar response profiles to these experimental factors. Third, eNODAL annotates these subclusters based on their experimental responses and biological pathways enriched within the subcluster. We validate eNODAL using data from a mouse experiment to test for the interaction effects of macronutrient intake and drugs that target aging mechanisms in mice.
Xiangnan Xu, Alistair M. Senior, David G. Le Couteur, Victoria C. Cogger, David Raubenheimer, David E. James, Stephen J. Simpson, Samuel Müller 0001, Jean Y. H. Yang
Briefings Bioinform.10
2025 CLUEY enables knowledge-guided clustering and cell type detection from single-cell omics data
abstract
MOTIVATION: Clustering is a fundamental task in single-cell omics data analysis and can significantly impact downstream analyses and biological interpretations. The standard approach involves grouping cells based on their gene expression profiles, followed by annotating each cluster to a cell type using marker genes. However, the number of cell types detected by different clustering methods can vary substantially due to several factors, including the dimension reduction method used and the choice of parameters of the chosen clustering algorithm. These discrepancies can lead to subjective interpretations in downstream analyses, particularly in manual cell type annotation. RESULTS: To address these challenges, we propose CLUEY, a knowledge-guided framework for cell type detection and clustering of single-cell omics data. CLUEY integrates prior biological knowledge into the clustering process, providing guidance on the optimal number of clusters and enhancing the interpretability of results. We apply CLUEY to both unimodal (e.g. scRNA-seq, scATAC-seq) and multimodal datasets (e.g. CITE-seq, SHARE-seq) and demonstrate its effectiveness in providing biologically meaningful clustering outcomes. These results highlight CLUEY on providing the much-needed guidance in clustering analyses of single-cell omics data. AVAILABILITY AND IMPLEMENTATION: CLUEY package is freely available from https://github.com/SydneyBioX/CLUEY.
Carissa Chen, Lijia Yu, Jean Y. H. Yang, Pengyi Yang
Bioinform.4
2025 IMPACT: interpretable microbial phenotype analysis via microbial characteristic traits
abstract
MOTIVATION: The human gut microbiome, consisting of trillions of bacteria, significantly impacts health and disease. High-throughput profiling through the advancement of modern technology provides the potential to enhance our understanding of the link between the microbiome and complex disease outcomes. However, there remains an open challenge where current microbiome models lack interpretability of microbial features, limiting a deeper understanding of the role of the gut microbiome in disease. To address this, we present a framework that combines a feature engineering step to transform tabular abundance data to image format using functional microbial annotation databases, with a residual spatial attention transformer block architecture for phenotype classification. RESULTS: Our model, IMPACT, delivers improved predictive accuracy performance across multiclass classification compared to similar methods. More importantly, our approach provides interpretable feature importance through image classification saliency methods. This enables the extraction of taxa markers (features) associated with a disease outcome and also their associated functional microbial traits and metabolites. AVAILABILITY AND IMPLEMENTATION: IMPACT is available at https://github.com/SydneyBioX/IMPACT. We providedirect installation of IMPACT via pip.
Daniel Mechtersheimer, Wenze Ding, Xiangnan Xu, Carolyn Sue, Jean Y. H. Yang
Bioinform.7
2025 Scope+: an open source generalizable architecture for single-cell RNA-seq atlases at sample and cell levels
abstract
SUMMARY: With the recent advancement in single-cell RNA-sequencing technologies and the increased availability of integrative tools, challenges arise in easy and fast access to large collections of cell atlas. Existing cell atlas portals rarely are open sourced and adaptable, and do not support meta-analysis at cell level. Here, we present an open source, highly optimized and scalable architecture, named Scope+, to allow quick access, meta-analysis and cell-level selection of the atlas data. We applied this architecture to our well-curated 5 million COVID-19 blood and immune cells, as a portal called Covidscope. We achieved efficient access to atlas-scale data via three strategies, such as cell-as-unit data modelling, novel database optimization techniques and innovative software architectural design. Scope+ serves as an open source architecture for researchers to build on with their own atlas. AVAILABILITY AND IMPLEMENTATION: The COVID-19 web portal, data and meta-analysis are available on Covidscope (https://covidsc.d24h.hk/). User tutorials on how to implement Scope+ architecture with their atlases can be found at https://hiyin.github.io/scopeplus-user-tutorial/. Scope+ source code can be found at https://doi.org/10.5281/zenodo.14174632 and https://github.com/hiyin/scopeplus.
Danqing Yin, Candice L. Y. Mak, Ken H. O. Yu, Yingxin Lin, Joshua W. K. Ho, Jean Y. H. Yang
Bioinform.10
2023 Benchmarking of analytical combinations for COVID-19 outcome prediction using single-cell RNA sequencing data
abstract
The advances of single-cell transcriptomic technologies have led to increasing use of single-cell RNA sequencing (scRNA-seq) data in large-scale patient cohort studies. The resulting high-dimensional data can be summarized and incorporated into patient outcome prediction models in several ways; however, there is a pressing need to understand the impact of analytical decisions on such model quality. In this study, we evaluate the impact of analytical choices on model choices, ensemble learning strategies and integrate approaches on patient outcome prediction using five scRNA-seq COVID-19 datasets. First, we examine the difference in performance between using single-view feature space versus multi-view feature space. Next, we survey multiple learning platforms from classical machine learning to modern deep learning methods. Lastly, we compare different integration approaches when combining datasets is necessary. Through benchmarking such analytical combinations, our study highlights the power of ensemble learning, consistency among different learning methods and robustness to dataset normalization when using multiple datasets as the model input.
Shila Ghazanfar, Pengyi Yang, Jean Y. H. Yang
Briefings Bioinform.4
2023 scSTAR reveals hidden heterogeneity with a real-virtual cell pair structure across conditions in single-cell RNA sequencing data
abstract
Cell-state transition can reveal additional information from single-cell ribonucleic acid (RNA)-sequencing data in time-resolved biological phenomena. However, most of the current methods are based on the time derivative of the gene expression state, which restricts them to the short-term evolution of cell states. Here, we present single-cell State Transition Across-samples of RNA-seq data (scSTAR), which overcomes this limitation by constructing a paired-cell projection between biological conditions with an arbitrary time span by maximizing the covariance between two feature spaces using partial least square and minimum squared error methods. In mouse ageing data, the response to stress in CD4+ memory T cell subtypes was found to be associated with ageing. A novel Treg subtype characterized by mTORC activation was identified to be associated with antitumour immune suppression, which was confirmed by immunofluorescence microscopy and survival analysis in 11 cancers from The Cancer Genome Atlas Program. On melanoma data, scSTAR improved immunotherapy-response prediction accuracy from 0.8 to 0.96.
Jiawei Zou, Duojiao Wu, Guoguo Shang, Jean Y. H. Yang, KongFatt Wong-Lin, Hourong Sun, Wantao Chen
Briefings Bioinform.8
2023 Ensemble deep learning of embeddings for clustering multimodal single-cell omics data
abstract
MOTIVATION: Recent advances in multimodal single-cell omics technologies enable multiple modalities of molecular attributes, such as gene expression, chromatin accessibility, and protein abundance, to be profiled simultaneously at a global level in individual cells. While the increasing availability of multiple data modalities is expected to provide a more accurate clustering and characterization of cells, the development of computational methods that are capable of extracting information embedded across data modalities is still in its infancy. RESULTS: We propose SnapCCESS for clustering cells by integrating data modalities in multimodal single-cell omics data using an unsupervised ensemble deep learning framework. By creating snapshots of embeddings of multimodality using variational autoencoders, SnapCCESS can be coupled with various clustering algorithms for generating consensus clustering of cells. We applied SnapCCESS with several clustering algorithms to various datasets generated from popular multimodal single-cell omics technologies. Our results demonstrate that SnapCCESS is effective and more efficient than conventional ensemble deep learning-based clustering methods and outperforms other state-of-the-art multimodal embedding generation methods in integrating data modalities for clustering cells. The improved clustering of cells from SnapCCESS will pave the way for more accurate characterization of cell identity and types, an essential step for various downstream analyses of multimodal single-cell omics data. AVAILABILITY AND IMPLEMENTATION: SnapCCESS is implemented as a Python package and is freely available from https://github.com/PYangLab/SnapCCESS under the open-source license of GPL-3. The data used in this study are publicly available (see section 'Data availability').
Lijia Yu, Jean Y. H. Yang, Pengyi Yang
Bioinform.3
2022 scFeatures: multi-view representations of single-cell and spatial data for disease outcome prediction
abstract
MOTIVATION: With the recent surge of large-cohort scale single cell research, it is of critical importance that analytical methods can fully utilize the comprehensive characterization of cellular systems that single cell technologies produce to provide insights into samples from individuals. Currently, there is little consensus on the best ways to compress information from the complex data structures of these technologies to summary statistics that represent each sample (e.g. individuals). RESULTS: Here, we present scFeatures, an approach that creates interpretable cellular and molecular representations of single-cell and spatial data at the sample level. We demonstrate that summarizing a broad collection of features at the sample level is both important for understanding underlying disease mechanisms in different experimental studies and for accurately classifying disease status of individuals. AVAILABILITY AND IMPLEMENTATION: scFeatures is publicly available as an R package at https://github.com/SydneyBioX/scFeatures. All data used in this study are publicly available with accession ID reported in the Section 2. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yingxin Lin, Ellis Patrick, Pengyi Yang, Jean Y. H. Yang
Bioinform.5
2022 Scalable workflow for characterization of cell-cell communication in COVID-19 patients
abstract
COVID-19 patients display a wide range of disease severity, ranging from asymptomatic to critical symptoms with high mortality risk. Our ability to understand the interaction of SARS-CoV-2 infected cells within the lung, and of protective or dysfunctional immune responses to the virus, is critical to effectively treat these patients. Currently, our understanding of cell-cell interactions across different disease states, and how such interactions may drive pathogenic outcomes, is incomplete. Here, we developed a generalizable and scalable workflow for identifying cells that are differentially interacting across COVID-19 patients with distinct disease outcomes and use this to examine eight public single-cell RNA-seq datasets (six from peripheral blood mononuclear cells, one from bronchoalveolar lavage and one from nasopharyngeal), with a total of 211 individual samples. By characterizing the cell-cell interaction patterns across epithelial and immune cells in lung tissues for patients with varying disease severity, we illustrate diverse communication patterns across individuals, and discover heterogeneous communication patterns among moderate and severe patients. We further illustrate patterns derived from cell-cell interactions are potential signatures for discriminating between moderate and severe patients. Overall, this workflow can be generalized and scaled to combine multiple scRNA-seq datasets to uncover cell-cell interactions.
Yingxin Lin, Lipin Loo, Andy Tran, David M. Lin, Cesar Moreno, Daniel Hesselson, G. Gregory Neely, Jean Y. H. Yang
PLoS Comput. Biol.8
2020 CiteFuse enables multi-modal analysis of CITE-seq data
abstract
MOTIVATION: Multi-modal profiling of single cells represents one of the latest technological advancements in molecular biology. Among various single-cell multi-modal strategies, cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq) allows simultaneous quantification of two distinct species: RNA and cell-surface proteins. Here, we introduce CiteFuse, a streamlined package consisting of a suite of tools for doublet detection, modality integration, clustering, differential RNA and protein expression analysis, antibody-derived tag evaluation, ligand-receptor interaction analysis and interactive web-based visualization of CITE-seq data. RESULTS: We demonstrate the capacity of CiteFuse to integrate the two data modalities and its relative advantage against data generated from single-modality profiling using both simulations and real-world CITE-seq data. Furthermore, we illustrate a novel doublet detection method based on a combined index of cell hashing and transcriptome data. Finally, we demonstrate CiteFuse for predicting ligand-receptor interactions by using multi-modal CITE-seq data. Collectively, we demonstrate the utility and effectiveness of CiteFuse for the integrative analysis of transcriptome and epitope profiles from CITE-seq data. AVAILABILITY AND IMPLEMENTATION: CiteFuse is freely available at http://shiny.maths.usyd.edu.au/CiteFuse/ as an online web service and at https://github.com/SydneyBioX/CiteFuse/ as an R package. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Hani Jieun Kim, Yingxin Lin, Thomas Andrew Geddes, Jean Y. H. Yang, Pengyi Yang
Bioinform.4
2020 LC-N2G: a local consistency approach for nutrigenomics data analysis
abstract
BACKGROUND: Nutrigenomics aims at understanding the interaction between nutrition and gene information. Due to the complex interactions of nutrients and genes, their relationship exhibits non-linearity. One of the most effective and efficient methods to explore their relationship is the nutritional geometry framework which fits a response surface for the gene expression over two prespecified nutrition variables. However, when the number of nutrients involved is large, it is challenging to find combinations of informative nutrients with respect to a certain gene and to test whether the relationship is stronger than chance. Methods for identifying informative combinations are essential to understanding the relationship between nutrients and genes. RESULTS: We introduce Local Consistency Nutrition to Graphics (LC-N2G), a novel approach for ranking and identifying combinations of nutrients with gene expression. In LC-N2G, we first propose a model-free quantity called Local Consistency statistic to measure whether there is non-random relationship between combinations of nutrients and gene expression measurements based on (1) the similarity between samples in the nutrient space and (2) their difference in gene expression. Then combinations with small LC are selected and a permutation test is performed to evaluate their significance. Finally, the response surfaces are generated for the subset of significant relationships. Evaluation on simulated data and real data shows the LC-N2G can accurately find combinations that are correlated with gene expression. CONCLUSION: The LC-N2G is practically powerful for identifying the informative nutrition variables correlated with gene expression. Therefore, LC-N2G is important in the area of nutrigenomics for understanding the relationship between nutrition and gene expression information.
Xiangnan Xu, Samantha Solon-Biet, Alistair M. Senior, David Raubenheimer, Stephen J. Simpson, Luigi Fontana, Samuel Müller 0001, Jean Y. H. Yang
BMC Bioinform.8
2019 Challenges for Brain Data Analysis in VR Environments
abstract
Analysing and understanding brain function and disorder is the main focus of neuroscience. Due to the high complexity of the brain, directionality of the signal and changing activity over time, visual exploration and data analysis are difficult. For this reason, a vast amount of research challenges are still unsolved. We explored different challenges of the visual analysis of brain data and the design of corresponding immersive environments in collaboration with experts from the biomedical domain. We built a prototype of an immersive virtual reality environment to explore the design space and to investigate how brain data analysis can be supported by a variety of design choices. Our environment can be used to study the effect of different visualisations and combinations of brain data representation, as for example network layouts, anatomical mapping or time series. As a long-term goal, we aim to aid neuro-scientists in a better understanding of brain function and disorder.
Sabrina Jaeger, Karsten Klein 0001, Lucas Joos, Johannes Zagermann, Michael de Ridder, Jinman Kim, Jean Y. H. Yang, Ulrike Pfeil, Harald Reiterer, Falk Schreiber
PacificVis7
2019 Impact of similarity metrics on single-cell RNA-seq data clustering
abstract
Advances in high-throughput sequencing on single-cell gene expressions [single-cell RNA sequencing (scRNA-seq)] have enabled transcriptome profiling on individual cells from complex samples. A common goal in scRNA-seq data analysis is to discover and characterise cell types, typically through clustering methods. The quality of the clustering therefore plays a critical role in biological discovery. While numerous clustering algorithms have been proposed for scRNA-seq data, fundamentally they all rely on a similarity metric for categorising individual cells. Although several studies have compared the performance of various clustering algorithms for scRNA-seq data, currently there is no benchmark of different similarity metrics and their influence on scRNA-seq data clustering. Here, we compared a panel of similarity metrics on clustering a collection of annotated scRNA-seq datasets. Within each dataset, a stratified subsampling procedure was applied and an array of evaluation measures was employed to assess the similarity metrics. This produced a highly reliable and reproducible consensus on their performance assessment. Overall, we found that correlation-based metrics (e.g. Pearson's correlation) outperformed distance-based metrics (e.g. Euclidean distance). To test if the use of correlation-based metrics can benefit the recently published clustering techniques for scRNA-seq data, we modified a state-of-the-art kernel-based clustering algorithm (SIMLR) using Pearson's correlation as a similarity measure and found significant performance improvement over Euclidean distance on scRNA-seq data clustering. These findings demonstrate the importance of similarity metrics in clustering scRNA-seq data and highlight Pearson's correlation as a favourable choice. Further comparison on different scRNA-seq library preparation protocols suggests that they may also affect clustering performance. Finally, the benchmarking framework is available at http://www.maths.usyd.edu.au/u/SMS/bioinformatics/software.html.
Taiyun Kim, Irene Rui Chen, Yingxin Lin, Andy Yi-Yang Wang, Jean Y. H. Yang, Pengyi Yang
Briefings Bioinform.5
2019 DCARS: differential correlation across ranked samples
abstract
MOTIVATION: Genes act as a system and not in isolation. Thus, it is important to consider coordinated changes of gene expression rather than single genes when investigating biological phenomena such as the aetiology of cancer. We have developed an approach for quantifying how changes in the association between pairs of genes may inform the outcome of interest called Differential Correlation across Ranked Samples (DCARS). Modelling gene correlation across a continuous sample ranking does not require the dichotomisation of samples into two distinct classes and can identify differences in gene correlation across early, mid or late stages of the outcome of interest. RESULTS: When we evaluated DCARS against the typical Fisher Z-transformation test for differential correlation, as well as a typical approach testing for interaction within a linear model, on real TCGA data, DCARS significantly ranked gene pairs containing known cancer genes more highly across several cancers. Similar results are found with our simulation study. DCARS was applied to 13 cancers datasets in TCGA, revealing several distinct relationships for which survival ranking was found to be associated with a change in correlation between genes. Furthermore, we demonstrated that DCARS can be used in conjunction with network analysis techniques to extract biological meaning from multi-layered and complex data. AVAILABILITY AND IMPLEMENTATION: DCARS R package and sample data are available at https://github.com/shazanfar/DCARS. Publicly available data from The Cancer Genome Atlas (TCGA) was used using the TCGABiolinks R package. Supplementary Files and DCARS R package is available at https://github.com/shazanfar/DCARS. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Shila Ghazanfar, Dario Strbenac, John T. Ormerod, Jean Y. H. Yang, Ellis Patrick
Bioinform.4
2019 bcGST - an interactive bias-correction method to identify over-represented gene-sets in boutique arrays
abstract
MOTIVATION: Gene annotation and pathway databases such as Gene Ontology and Kyoto Encyclopaedia of Genes and Genomes are important tools in Gene-Set Test (GST) that describe gene biological functions and associated pathways. GST aims to establish an association relationship between a gene-set of interest and an annotation. Importantly, GST tests for over-representation of genes in an annotation term. One implicit assumption of GST is that the gene expression platform captures the complete or a very large proportion of the genome. However, this assumption is neither satisfied for the increasingly popular boutique array nor the custom designed gene expression profiling platform. Specifically, conventional GST is no longer appropriate due to the gene-set selection bias induced during the construction of these platforms. RESULTS: We propose bcGST, a bias-corrected GST by introducing bias-correction terms in the contingency table needed for calculating the Fisher's Exact Test. The adjustment method works by estimating the proportion of genes captured on the array with respect to the genome in order to assist filtration of annotation terms that would otherwise be falsely included or excluded. We illustrate the practicality of bcGST and its stability through multiple differential gene expression analyses in melanoma and the Cancer Genome Atlas cancer studies. AVAILABILITY AND IMPLEMENTATION: The bcGST method is made available as a Shiny web application at http://shiny.maths.usyd.edu.au/bcGST/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Kevin Y. X. Wang, Alexander M. Menzies, Ines P. Silva, James S. Wilmott, Yibing Yan, Matthew Wongchenko, Richard F. Kefford, Richard A. Scolyer, Georgina V. Long, Garth Tarr, Samuel Müller 0001, Jean Y. H. Yang
Bioinform.12
2019 scDC: single cell differential composition analysis
abstract
BACKGROUND: Differences in cell-type composition across subjects and conditions often carry biological significance. Recent advancements in single cell sequencing technologies enable cell-types to be identified at the single cell level, and as a result, cell-type composition of tissues can now be studied in exquisite detail. However, a number of challenges remain with cell-type composition analysis - none of the existing methods can identify cell-type perfectly and variability related to cell sampling exists in any single cell experiment. This necessitates the development of method for estimating uncertainty in cell-type composition. RESULTS: We developed a novel single cell differential composition (scDC) analysis method that performs differential cell-type composition analysis via bootstrap resampling. scDC captures the uncertainty associated with cell-type proportions of each subject via bias-corrected and accelerated bootstrap confidence intervals. We assessed the performance of our method using a number of simulated datasets and synthetic datasets curated from publicly available single cell datasets. In simulated datasets, scDC correctly recovered the true cell-type proportions. In synthetic datasets, the cell-type compositions returned by scDC were highly concordant with reference cell-type compositions from the original data. Since the majority of datasets tested in this study have only 2 to 5 subjects per condition, the addition of confidence intervals enabled better comparisons of compositional differences between subjects and across conditions. CONCLUSIONS: scDC is a novel statistical method for performing differential cell-type composition analysis for scRNA-seq data. It uses bootstrap resampling to estimate the standard errors associated with cell-type proportion estimates and performs significance testing through GLM and GLMM models. We have made this method available to the scientific community as part of the scdney package (Single Cell Data Integrative Analysis) R package, available from https://github.com/SydneyBioX/scdney.
Yingxin Lin, John T. Ormerod, Pengyi Yang, Jean Y. H. Yang, Kitty K. Lo
BMC Bioinform.5
2019 Autoencoder-based cluster ensembles for single-cell RNA-seq data analysis
abstract
BACKGROUND: Single-cell RNA-sequencing (scRNA-seq) is a transformative technology, allowing global transcriptomes of individual cells to be profiled with high accuracy. An essential task in scRNA-seq data analysis is the identification of cell types from complex samples or tissues profiled in an experiment. To this end, clustering has become a key computational technique for grouping cells based on their transcriptome profiles, enabling subsequent cell type identification from each cluster of cells. Due to the high feature-dimensionality of the transcriptome (i.e. the large number of measured genes in each cell) and because only a small fraction of genes are cell type-specific and therefore informative for generating cell type-specific clusters, clustering directly on the original feature/gene dimension may lead to uninformative clusters and hinder correct cell type identification. RESULTS: Here, we propose an autoencoder-based cluster ensemble framework in which we first take random subspace projections from the data, then compress each random projection to a low-dimensional space using an autoencoder artificial neural network, and finally apply ensemble clustering across all encoded datasets to generate clusters of cells. We employ four evaluation metrics to benchmark clustering performance and our experiments demonstrate that the proposed autoencoder-based cluster ensemble can lead to substantially improved cell type-specific clusters when applied with both the standard k-means clustering algorithm and a state-of-the-art kernel-based clustering algorithm (SIMLR) designed specifically for scRNA-seq data. Compared to directly using these clustering algorithms on the original datasets, the performance improvement in some cases is up to 100%, depending on the evaluation metric used. CONCLUSIONS: Our results suggest that the proposed framework can facilitate more accurate cell type identification as well as other downstream analyses. The code for creating the proposed autoencoder-based cluster ensemble framework is freely available from https://github.com/gedcom/scCCESS.
Thomas Andrew Geddes, Taiyun Kim, Lihao Nan, James G. Burchfield, Jean Y. H. Yang, Dacheng Tao, Pengyi Yang
BMC Bioinform.5
2019 AdaSampling for Positive-Unlabeled and Label Noise Learning With Bioinformatics Applications
abstract
Class labels are required for supervised learning but may be corrupted or missing in various applications. In binary classification, for example, when only a subset of positive instances is labeled whereas the remaining are unlabeled, positive-unlabeled (PU) learning is required to model from both positive and unlabeled data. Similarly, when class labels are corrupted by mislabeled instances, methods are needed for learning in the presence of class label noise (LN). Here we propose adaptive sampling (AdaSampling), a framework for both PU learning and learning with class LN. By iteratively estimating the class mislabeling probability with an adaptive sampling procedure, the proposed method progressively reduces the risk of selecting mislabeled instances for model training and subsequently constructs highly generalizable models even when a large proportion of mislabeled instances is present in the data. We demonstrate the utilities of proposed methods using simulation and benchmark data, and compare them to alternative approaches that are commonly used for PU learning and/or learning with LN. We then introduce two novel bioinformatics applications where AdaSampling is used to: 1) identify kinase-substrates from mass spectrometry-based phosphoproteomics data and 2) predict transcription factor target genes by integrating various next-generation sequencing data.
Pengyi Yang, John T. Ormerod, Wei Liu 0007, Chendong Ma, Albert Y. Zomaya, Jean Y. H. Yang
IEEE Trans. Cybern.6
2017 Positive unlabeled learning via wrapper-based adaptive sampling
abstract
Learning from positive and unlabeled data frequently occurs in applications where only a subset of positive instances is available while the rest of the data are unlabeled. In such scenarios, often the goal is to create a discriminant model that can accurately classify both positive and negative data by modelling from labeled and unlabeled instances. In this study, we propose an adaptive sampling (AdaSampling) approach that utilises prediction probabilities from a model to iteratively update the training data. Starting with equal prior probabilities for all unlabeled data, our method "wraps" around a predictive model to iteratively update these probabilities to distinguish positive and negative instances in unlabeled data. Subsequently, one or more robust negative set(s) can be drawn from unlabeled data, according to the likelihood of each instance being negative, to train a single classification model or ensemble of models.
Pengyi Yang, Wei Liu 0007, Jean Y. H. Yang
IJCAI3
2017 Integrative analysis identifies co-dependent gene expression regulation of BRG1 and CHD7 at distal regulatory sites in embryonic stem cells
abstract
MOTIVATION: DNA binding proteins such as chromatin remodellers, transcription factors (TFs), histone modifiers and co-factors often bind cooperatively to activate or repress their target genes in a cell type-specific manner. Nonetheless, the precise role of cooperative binding in defining cell-type identity is still largely uncharacterized. RESULTS: Here, we collected and analyzed 214 public datasets representing chromatin immunoprecipitation followed by sequencing (ChIP-Seq) of 104 DNA binding proteins in embryonic stem cell (ESC) lines. We classified their binding sites into those proximal to gene promoters and those in distal regions, and developed a web resource called Proximal And Distal (PAD) clustering to identify their co-localization at these respective regions. Using this extensive dataset, we discovered an extensive co-localization of BRG1 and CHD7 at distal but not proximal regions. The comparison of co-localization sites to those bound by either BRG1 or CHD7 alone showed an enrichment of ESC master TFs binding and active chromatin architecture at co-localization sites. Most notably, our analysis reveals the co-dependency of BRG1 and CHD7 at distal regions on regulating expression of their common target genes in ESC. This work sheds light on cooperative binding of TF binding proteins in regulating gene expression in ESC, and demonstrates the utility of integrative analysis of a manually curated compendium of genome-wide protein binding profiles in our online resource PAD. AVAILABILITY AND IMPLEMENTATION: PAD is freely available at http://pad.victorchang.edu.au/ and its source code is available via an open source GPL 3.0 license at https://github.com/VCCRI/PAD/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengyi Yang, Andrew J. Oldfield, Taiyun Kim, Andrian Yang, Jean Y. H. Yang, Joshua W. K. Ho
Bioinform.5
2016 Positive-unlabeled ensemble learning for kinase substrate prediction from dynamic phosphoproteomics data
abstract
MOTIVATION: Protein phosphorylation is a post-translational modification that underlines various aspects of cellular signaling. A key step to reconstructing signaling networks involves identification of the set of all kinases and their substrates. Experimental characterization of kinase substrates is both expensive and time-consuming. To expedite the discovery of novel substrates, computational approaches based on kinase recognition sequence (motifs) from known substrates, protein structure, interaction and co-localization have been proposed. However, rarely do these methods take into account the dynamic responses of signaling cascades measured from in vivo cellular systems. Given that recent advances in mass spectrometry-based technologies make it possible to quantify phosphorylation on a proteome-wide scale, computational approaches that can integrate static features with dynamic phosphoproteome data would greatly facilitate the prediction of biologically relevant kinase-specific substrates. RESULTS: Here, we propose a positive-unlabeled ensemble learning approach that integrates dynamic phosphoproteomics data with static kinase recognition motifs to predict novel substrates for kinases of interest. We extended a positive-unlabeled learning technique for an ensemble model, which significantly improves prediction sensitivity on novel substrates of kinases while retaining high specificity. We evaluated the performance of the proposed model using simulation studies and subsequently applied it to predict novel substrates of key kinases relevant to insulin signaling. Our analyses show that static sequence motifs and dynamic phosphoproteomics data are complementary and that the proposed integrated model performs better than methods relying only on static information for accurate prediction of kinase-specific substrates. AVAILABILITY AND IMPLEMENTATION: Executable GUI tool, source code and documentation are freely available at https://github.com/PengyiYang/KSP-PUEL. CONTACT: [email protected] or [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengyi Yang, Sean J. Humphrey, David E. James, Jean Y. H. Yang, Raja Jothi
Bioinform.4
2015 Inferring data-specific micro-RNA function through the joint ranking of micro-RNA and pathways from matched micro-RNA and gene expression data
abstract
MOTIVATION: In practice, identifying and interpreting the functional impacts of the regulatory relationships between micro-RNA and messenger-RNA is non-trivial. The sheer scale of possible micro-RNA and messenger-RNA interactions can make the interpretation of results difficult. RESULTS: We propose a supervised framework, pMim, built upon concepts of significance combination, for jointly ranking regulatory micro-RNA and their potential functional impacts with respect to a condition of interest. Here, pMim directly tests if a micro-RNA is differentially expressed and if its predicted targets, which lie in a common biological pathway, have changed in the opposite direction. We leverage the information within existing micro-RNA target and pathway databases to stabilize the estimation and annotation of micro-RNA regulation making our approach suitable for datasets with small sample sizes. In addition to outputting meaningful and interpretable results, we demonstrate in a variety of datasets that the micro-RNA identified by pMim, in comparison to simpler existing approaches, are also more concordant with what is described in the literature. AVAILABILITY AND IMPLEMENTATION: This framework is implemented as an R function, pMim, in the package sydSeq available from http://www.ellispatrick.com/r-packages. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Ellis Patrick, Michael J. Buckley, Samuel Müller 0001, David M. Lin, Jean Y. H. Yang
Bioinform.5
2015 ClassifyR: an R package for performance assessment of classification with applications to transcriptomics
abstract
UNLABELLED: Although a large collection of classification software packages exist in R, a new generic framework for linking custom classification functions with classification performance measures is needed. A generic classification framework has been designed and implemented as an R package in an object oriented style. Its design places emphasis on parallel processing, reproducibility and extensibility. Finally, a comprehensive set of performance measures are available to ease post-processing. Taken together, these important characteristics enable rapid and reproducible benchmarking of alternative classifiers. AVAILABILITY AND IMPLEMENTATION: ClassifyR is implemented in R and can be obtained from the Bioconductor project: http://bioconductor.org/packages/release/bioc/html/ClassifyR.html.
Dario Strbenac, Graham J. Mann, John T. Ormerod, Jean Y. H. Yang
Bioinform.4
2015 Knowledge-Based Analysis for Detecting Key Signaling Events from Time-Series Phosphoproteomics Data
abstract
Cell signaling underlies transcription/epigenetic control of a vast majority of cell-fate decisions. A key goal in cell signaling studies is to identify the set of kinases that underlie key signaling events. In a typical phosphoproteomics study, phosphorylation sites (substrates) of active kinases are quantified proteome-wide. By analyzing the activities of phosphorylation sites over a time-course, the temporal dynamics of signaling cascades can be elucidated. Since many substrates of a given kinase have similar temporal kinetics, clustering phosphorylation sites into distinctive clusters can facilitate identification of their respective kinases. Here we present a knowledge-based CLUster Evaluation (CLUE) approach for identifying the most informative partitioning of a given temporal phosphoproteomics data. Our approach utilizes prior knowledge, annotated kinase-substrate relationships mined from literature and curated databases, to first generate biologically meaningful partitioning of the phosphorylation sites and then determine key kinases associated with each cluster. We demonstrate the utility of the proposed approach on two time-series phosphoproteomics datasets and identify key kinases associated with human embryonic stem cell differentiation and insulin signaling pathway. The proposed approach will be a valuable resource in the identification and characterizing of signaling networks from phosphoproteomics data.
Pengyi Yang, Vivek Jayaswal, Guang Hu 0005, Jean Y. H. Yang, Raja Jothi
PLoS Comput. Biol.5
2014 Direction pathway analysis of large-scale proteomics data reveals novel features of the insulin action pathway
abstract
MOTIVATION: With the advancement of high-throughput techniques, large-scale profiling of biological systems with multiple experimental perturbations is becoming more prevalent. Pathway analysis incorporates prior biological knowledge to analyze genes/proteins in groups in a biological context. However, the hypotheses under investigation are often confined to a 1D space (i.e. up, down, either or mixed regulation). Here, we develop direction pathway analysis (DPA), which can be applied to test hypothesis in a high-dimensional space for identifying pathways that display distinct responses across multiple perturbations. RESULTS: Our DPA approach allows for the identification of pathways that display distinct responses across multiple perturbations. To demonstrate the utility and effectiveness, we evaluated DPA under various simulated scenarios and applied it to study insulin action in adipocytes. A major action of insulin in adipocytes is to regulate the movement of proteins from the interior to the cell surface membrane. Quantitative mass spectrometry-based proteomics was used to study this process on a large-scale. The combined dataset comprises four separate treatments. By applying DPA, we identified that several insulin responsive pathways in the plasma membrane trafficking are only partially dependent on the insulin-regulated kinase Akt. We subsequently validated our findings through targeted analysis of key proteins from these pathways using immunoblotting and live cell microscopy. Our results demonstrate that DPA can be applied to dissect pathway networks testing diverse hypotheses and integrating multiple experimental perturbations. AVAILABILITY AND IMPLEMENTATION: The R package 'directPA' is distributed from CRAN under GNU General Public License (GPL)-3 and can be downloaded from: http://cran.r-project.org/web/packages/directPA/index.html CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Pengyi Yang, Ellis Patrick, Shi-Xiong Tan, Daniel J. Fazakerley, James G. Burchfield, Christopher Gribben, Matthew J. Prior, David E. James, Jean Y. H. Yang
Bioinform.9
2013 Latin square dataset for evaluating the accuracy of mass spectrometry-based protein identification and quantification
abstract
Tandem mass spectrometry-based iTRAQ protein quantification provides a powerful means for identifying disease biomarkers and plays an important role in developing new diagnosis and prognosis, new treatment, and personalised medicine. However, analyses of iTRAQ data encounter a number of statistical and computational challenges such as accurate protein identification, imputation of missing values, appropriate summarisation of protein quantification, among others. Therefore, a good evaluation dataset where raw spectra are provided, and actual composition and concentrations of the protein mixture are known, will enable better methodological development in this field. Unfortunately, there are limited evaluation datasets and existing ones are not sufficient for systematic evaluation of the existing analysis methods. To this end, we designed and performed a new Latin Square experiment that can be used for validating the accuracy of both protein identification and protein quantification.
Penghao Wang 0002, Jean Y. H. Yang, Mark Raftery, Susan R. Wilson
BIBM2
2013 Estimation of data-specific constitutive exons with RNA-Seq data
abstract
BACKGROUND: RNA-Seq has the potential to answer many diverse and interesting questions about the inner workings of cells. Estimating changes in the overall transcription of a gene is not straightforward. Changes in overall gene transcription can easily be confounded with changes in exon usage which alter the lengths of transcripts produced by a gene. Measuring the expression of constitutive exons--xons which are consistently conserved after splicing--ffers an unbiased estimation of the overall transcription of a gene. RESULTS: We propose a clustering-based method, exClust, for estimating the exons that are consistently conserved after splicing in a given data set. These are considered as the exons which are "constitutive" in this data. The method utilises information from both annotation and the dataset of interest. The method is implemented in an openly available R function package, sydSeq. CONCLUSION: When used on two real datasets exClust includes more than three times as many reads as the standard UI method, and improves concordance with qRT-PCR data. When compared to other methods, our method is shown to produce robust estimates of overall gene transcription.
Ellis Patrick, Michael J. Buckley, Jean Y. H. Yang
BMC Bioinform.3
2012 OCAP: an open comprehensive analysis pipeline for iTRAQ
abstract
MOTIVATION: Mass spectrometry-based iTRAQ protein quantification is a high-throughput assay for determining relative protein expressions and identifying disease biomarkers. Processing and analysis of these large and complex data involves a number of distinct components and it is desirable to have a pipeline to efficiently integrate these together. To date, there are limited public available comprehensive analysis pipelines for iTRAQ data and many of these existing pipelines have limited visualization tools and no convenient interfaces with downstream analyses. We have developed a new open source comprehensive iTRAQ analysis pipeline, OCAP, integrating a wavelet-based preprocessing algorithm which provides better peak picking, a new quantification algorithm and a suite of visualizsation tools. OCAP is mainly developed in C++ and is provided as a standalone version (OCAP_standalone) as well as an R package. The R package (OCAP) provides the necessary interfaces with downstream statistical analysis.
Penghao Wang 0002, Pengyi Yang, Jean Y. H. Yang
Bioinform.3
2012 Improving X!Tandem on Peptide Identification from Mass Spectrometry by Self-Boosted Percolator
abstract
A critical component in mass spectrometry (MS)-based proteomics is an accurate protein identification procedure. Database search algorithms commonly generate a list of peptide-spectrum matches (PSMs). The validity of these PSMs is critical for downstream analysis since proteins that are present in the sample are inferred from those PSMs. A variety of postprocessing algorithms have been proposed to validate and filter PSMs. Among them, the most popular ones include a semi-supervised learning (SSL) approach known as Percolator and an empirical modeling approach known as PeptideProphet. However, they are predominantly designed for commercial database search algorithms, i.e., SEQUEST and MASCOT. Therefore, it is highly desirable to extend and optimize those PSM postprocessing algorithms for open source database search algorithms such as X!Tandem. In this paper, we propose a Self-boosted Percolator for postprocessing X!Tandem search results. We find that the SSL algorithm utilized by Percolator depends heavily on the initial ranking of PSMs. Starting with a poor PSM ranking list may cause Percolator to perform suboptimally. By implementing Percolator in a cascade learning manner, we can progressively improve the performance through multiple boost runs, enabling many more PSM identifications without sacrificing false discovery rate (FDR).
Pengyi Yang, Penghao Wang 0002, Bing Bing Zhou, Jean Y. H. Yang
IEEE ACM Trans. Comput. Biol. Bioinform.6
2011 Gene-gene interaction filtering with ensemble of filters
abstract
BACKGROUND: Complex diseases are commonly caused by multiple genes and their interactions with each other. Genome-wide association (GWA) studies provide us the opportunity to capture those disease associated genes and gene-gene interactions through panels of SNP markers. However, a proper filtering procedure is critical to reduce the search space prior to the computationally intensive gene-gene interaction identification step. In this study, we show that two commonly used SNP-SNP interaction filtering algorithms, ReliefF and tuned ReliefF (TuRF), are sensitive to the order of the samples in the dataset, giving rise to unstable and suboptimal results. However, we observe that the 'unstable' results from multiple runs of these algorithms can provide valuable information about the dataset. We therefore hypothesize that aggregating results from multiple runs of the algorithm may improve the filtering performance. RESULTS: We propose a simple and effective ensemble approach in which the results from multiple runs of an unstable filter are aggregated based on the general theory of ensemble learning. The ensemble versions of the ReliefF and TuRF algorithms, referred to as ReliefF-E and TuRF-E, are robust to sample order dependency and enable a more informative investigation of data characteristics. Using simulated and real datasets, we demonstrate that both the ensemble of ReliefF and the ensemble of TuRF can generate a much more stable SNP ranking than the original algorithms. Furthermore, the ensemble of TuRF achieved the highest success rate in comparison to many state-of-the-art algorithms as well as traditional χ2-test and odds ratio methods in terms of retaining gene-gene interactions.
Pengyi Yang, Joshua W. K. Ho, Jean Y. H. Yang, Bing Bing Zhou
BMC Bioinform.3
2011 Two-Step Cross-Entropy Feature Selection for Microarrays - Power Through Complementarity
abstract
Current feature selection methods for supervised classification of tissue samples from microarray data generally fail to exploit complementary discriminatory power that can be found in sets of features. Using a feature selection method with the computational architecture of the cross-entropy method, including an additional preliminary step ensuring a lower bound on the number of times any feature is considered, we show when testing on a human lymph node data set that there are a significant number of genes that perform well when their complementary power is assessed, but “pass under the radar” of popular feature selection methods that only assess genes individually on a given classification tool. We also show that this phenomenon becomes more apparent as diagnostic specificity of the tissue samples analysed increases.
Tim Peters, David W. Bulger, To-ha Loi, Jean Y. H. Yang, David Ma
IEEE ACM Trans. Comput. Biol. Bioinform.4
2010 A dynamic wavelet-based algorithm for pre-processing tandem mass spectrometry data
abstract
MOTIVATION: Mass spectrometry (MS)-based proteomics is one of the most commonly used research techniques for identifying and characterizing proteins in biological and medical research. The identification of a protein is the critical first step in elucidating its biological function. Successful protein identification depends on various interrelated factors, including effective analysis of MS data generated in a proteomic experiment. This analysis comprises several stages, often combined in a pipeline or workflow. The first component of the analysis is known as spectra pre-processing. In this component, the raw data generated by the mass spectrometer is processed to eliminate noise and identify the mass-to-charge ratio (m/z) and intensity for the peaks in the spectrum corresponding to the presence of certain peptides or peptide fragments. Since all downstream analyses depend on the pre-processed data, effective pre-processing is critical to protein identification and characterization. There is a critical need for more robust pre-processing algorithms that perform well on tandem mass spectra under a variety of different conditions and can be easily integrated into sophisticated data analysis pipelines for practical wet-lab applications. RESULT: We have developed a new pre-processing algorithm. Based on wavelet theory, our method uses a dynamic peak model to identify peaks. It is designed to be easily integrated into a complete proteomic analysis workflow. We compared the method with other available algorithms using a reference library of raw MS and tandem MS spectra with known protein composition information. Our pre-processing algorithm results in the identification of significantly more peptides and proteins in the downstream analysis for a given false discovery rate. AVAILABILITY: Software available at: http://www.maths.usyd.edu.au/u/penghao/index.html.
Penghao Wang 0002, Pengyi Yang, Jonathan Arthur, Jean Y. H. Yang
Bioinform.4
2010 Comparison study of microarray meta-analysis methods
abstract
BACKGROUND: Meta-analysis methods exist for combining multiple microarray datasets. However, there are a wide range of issues associated with microarray meta-analysis and a limited ability to compare the performance of different meta-analysis methods. RESULTS: We compare eight meta-analysis methods, five existing methods, two naive methods and a novel approach (mDEDS). Comparisons are performed using simulated data and two biological case studies with varying degrees of meta-analysis complexity. The performance of meta-analysis methods is assessed via ROC curves and prediction accuracy where applicable. CONCLUSIONS: Existing meta-analysis methods vary in their ability to perform successful meta-analysis. This success is very dependent on the complexity of the data and type of analysis. Our proposed method, mDEDS, performs competitively as a meta-analysis tool even as complexity increases. Because of the varying abilities of compared meta-analysis methods, care should be taken when considering the meta-analysis method used for particular research.
Anna Campain, Jean Y. H. Yang
BMC Bioinform.2
2007 A multi-array multi-SNP genotyping algorithm for Affymetrix SNP microarrays
abstract
MOTIVATION: Modern strategies for mapping disease loci require efficient genotyping of a large number of known polymorphic sites in the genome. The sensitive and high-throughput nature of hybridization-based DNA microarray technology provides an ideal platform for such an application by interrogating up to hundreds of thousands of single nucleotide polymorphisms (SNPs) in a single assay. Similar to the development of expression arrays, these genotyping arrays pose many data analytic challenges that are often platform specific. Affymetrix SNP arrays, e.g. use multiple sets of short oligonucleotide probes for each known SNP, and require effective statistical methods to combine these probe intensities in order to generate reliable and accurate genotype calls. RESULTS: We developed an integrated multi-SNP, multi-array genotype calling algorithm for Affymetrix SNP arrays, MAMS, that combines single-array multi-SNP (SAMS) and multi-array, single-SNP (MASS) calls to improve the accuracy of genotype calls, without the need for training data or computation-intensive normalization procedures as in other multi-array methods. The algorithm uses resampling techniques and model-based clustering to derive single array based genotype calls, which are subsequently refined by competitive genotype calls based on (MASS) clustering. The resampling scheme caps computation for single-array analysis and hence is readily scalable, important in view of expanding numbers of SNPs per array. The MASS update is designed to improve calls for atypical SNPs, harboring allele-imbalanced binding affinities, that are difficult to genotype without information from other arrays. Using a publicly available data set of HapMap samples from Affymetrix, and independent calls by alternative genotyping methods from the HapMap project, we show that our approach performs competitively to existing methods. AVAILABILITY: R functions are available upon request from the authors.
Yuanyuan Xiao, Mark R. Segal, Jean Y. H. Yang, Ru-Fang Yeh
Bioinform.3
2005 Identifying differentially expressed genes from microarray experiments via statistic synthesis
abstract
Abstract Motivation: A common objective of microarray experiments is the detection of differential gene expression between samples obtained under different conditions. The task of identifying differentially expressed genes consists of two aspects: ranking and selection. Numerous statistics have been proposed to rank genes in order of evidence for differential expression. However, no one statistic is universally optimal and there is seldom any basis or guidance that can direct toward a particular statistic of choice. Results: Our new approach, which addresses both ranking and selection of differentially expressed genes, integrates differing statistics via a distance synthesis scheme. Using a set of (Affymetrix) spike-in datasets, in which differentially expressed genes are known, we demonstrate that our method compares favorably with the best individual statistics, while achieving robustness properties lacked by the individual statistics. We further evaluate performance on one other microarray study. Availability: The approach is implemented in an R package called DEDS, which is available for download from the Bioconductor website (http://www.bioconductor.org/). Contact: [email protected]
Jean Y. H. Yang, Yuanyuan Xiao, Mark R. Segal
Bioinform.1
2005 Analysis of a Splice Array Experiment Elucidates Roles of Chromatin Elongation Factor Spt4-5 in Splicing
abstract
Splicing is an important process for regulation of gene expression in eukaryotes, and it has important functional links to other steps of gene expression. Two examples of these linkages include Ceg1, a component of the mRNA capping enzyme, and the chromatin elongation factors Spt4-5, both of which have recently been shown to play a role in the normal splicing of several genes in the yeast Saccharomyces cerevisiae. Using a genomic approach to characterize the roles of Spt4-5 in splicing, we used splicing-sensitive DNA microarrays to identify specific sets of genes that are mis-spliced in ceg1, spt4, and spt5 mutants. In the context of a complex, nested, experimental design featuring 22 dye-swap array hybridizations, comprising both biological and technical replicates, we applied five appropriate statistical models for assessing differential expression between wild-type and the mutants. To refine selection of differential expression genes, we then used a robust model-synthesizing approach, Differential Expression via Distance Synthesis, to integrate all five models. The resultant list of differentially expressed genes was then further analyzed with regard to select attributes: we found that highly transcribed genes with long introns were most sensitive to spt mutations. QPCR confirmation of differential expression was established for the limited number of genes evaluated. In this paper, we showcase splicing array technology, as well as powerful, yet general, statistical methodology for assessing differential expression, in the context of a real, complex experimental design. Our results suggest that the Spt4-Spt5 complex may help coordinate splicing with transcription under conditions that present kinetic challenges to spliceosome assembly or function.
Yuanyuan Xiao, Jean Y. H. Yang, Todd A. Burckin, Lily Shiue, Grant A. Hartzog, Mark R. Segal
PLoS Comput. Biol.2
2001 Analysis of CDNA Microarray Images
abstract
Microarrays are part of a new class of biotechnologies that allow the monitoring of expression levels for thousands of genes simultaneously. Image analysis is an important aspect of microarray experiments, one that can have a potentially large impact on subsequent analyses, such as clustering or the identification of differentially expressed genes. This paper reviews a number of existing image analysis methods used on cDNA microarray data. In particular, it describes and discusses the different segmentation and background adjustment methods. It was found that in some cases background adjustment can substantially reduce the precision--that is, increase the variability of low-intensity spot values. In contrast, the choice of segmentation procedure seems to have a smaller impact.
Jean Y. H. Yang, Michael J. Buckley, Terence P. Speed
Briefings Bioinform.1