EDBT 2026 Demo / reviewers in the wild / expert
Joshua W. K. Ho
dblp:78/5950 · also Joshua Wing Kei Ho
· DBLP profile ↗
23ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0003-2331-7011ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 20 · 2 first-author · 7 since 2021Theory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Should we use large language models to perform systematic reviews?abstractAbstract Background Large language models (LLMs) have shown promising performances in literature search, screening of full-text scientific articles, data extraction, and produce a written summary. LLMs have a strong potential in automating the process of conducting systematic reviews. Nonetheless, conducting systematic reviews is also a training exercise for critical appraisal of scientific evidence and synthesis of scientific claims. Overreliance on LLMs may erode comprehension and analytical skills in reading scientific articles over time, leading to cognitive offloading. Other risks may involve hallucination and responsibility problem. Methods To address both opportunities and limitations of LLMs when applied in systematic reviews, the LLM-Augmented Systematic Evidence Review (LASER) framework is proposed. Results The framework recommends different degree of LLM usage in different steps in the systematic review process to promote efficiency, ensure accuracy, and prevent cognitive offloading of the researcher involved. In this framework, the specific degree of LLM usage may depend on the career stage of the researcher and the accuracy of the LLM. Th Chan, Sh Jin, Mun Kay Ho, Joshua W. K. Ho |
Briefings Bioinform. | 4 |
| 2025 | Discovery of cell-type-specific mtDNA common deletion in nasopharyngeal carcinomaabstractAbstract Background The human mitochondrial genome (mtDNA) plays a key role in maintaining cellular functions. However, mtDNA is vulnerable to mutations due to the lack of protective histones and low efficiency in DNA repair mechanisms. Among these mutations, the 4977 bp mtDNA common deletion (mtDNA-CD) encompasses twelve essential genes that are essential for oxidative phosphorylation. Aim In previous studies, mtDNA-CD was identified in nasopharyngeal carcinoma (NPC) patient samples. However, bulk sequencing protocols employed in past studies were unable to distinguish whether those mtDNA-CD were specific to the cancer cells or from other non-cancer cell types in the samples. Single-cell RNA sequencing (scRNA-seq) enables the discovery of different cell types within a population and facilitates the identification of cell-type-specific mutation signatures. In this study, we propose to use scRNA-seq data to identify cell-type-specific deletion signatures in NPC samples. Methods We developed a custom computational pipeline to detect large-scale deletions using scRNA-seq data. The pipeline harnesses both clipped read profiles and secondary alignments to overcome the typical limitations of low sequencing depth and uneven read coverage in scRNA-seq data. We tested our pipeline on an NPC data set consisting of nine patients. Results Our result suggests the presence of mtDNA-CD mutations specific to malignant cancer cells in NPC tumour samples. This study proposes the use of scRNA-seq of cell-type-specific mtDNA-CD in NPC samples, offering new insights into the genetic and molecular pathogenesis of cancer and guiding the development of specific biomarkers for NPC. Mun Kay Ho, Joshua W. K. Ho |
Briefings Bioinform. | 2 |
| 2025 | Scope+: an open source generalizable architecture for single-cell RNA-seq atlases at sample and cell levelsabstractSUMMARY: With the recent advancement in single-cell RNA-sequencing technologies and the increased availability of integrative tools, challenges arise in easy and fast access to large collections of cell atlas. Existing cell atlas portals rarely are open sourced and adaptable, and do not support meta-analysis at cell level. Here, we present an open source, highly optimized and scalable architecture, named Scope+, to allow quick access, meta-analysis and cell-level selection of the atlas data. We applied this architecture to our well-curated 5 million COVID-19 blood and immune cells, as a portal called Covidscope. We achieved efficient access to atlas-scale data via three strategies, such as cell-as-unit data modelling, novel database optimization techniques and innovative software architectural design. Scope+ serves as an open source architecture for researchers to build on with their own atlas. AVAILABILITY AND IMPLEMENTATION: The COVID-19 web portal, data and meta-analysis are available on Covidscope (https://covidsc.d24h.hk/). User tutorials on how to implement Scope+ architecture with their atlases can be found at https://hiyin.github.io/scopeplus-user-tutorial/. Scope+ source code can be found at https://doi.org/10.5281/zenodo.14174632 and https://github.com/hiyin/scopeplus. Danqing Yin, Candice L. Y. Mak, Ken H. O. Yu, Yingxin Lin, Joshua W. K. Ho, Jean Y. H. Yang |
Bioinform. | 9 |
| 2024 | scDecouple: decoupling cellular response from infected proportion bias in scCRISPR-seqabstractSingle-cell clustered regularly interspaced short palindromic repeats-sequencing (scCRISPR-seq) is an emerging high-throughput CRISPR screening technology where the true cellular response to perturbation is coupled with infected proportion bias of guide RNAs (gRNAs) across different cell clusters. The mixing of these effects introduces noise into scCRISPR-seq data analysis and thus obstacles to relevant studies. We developed scDecouple to decouple true cellular response of perturbation from the influence of infected proportion bias. scDecouple first models the distribution of gene expression profiles in perturbed cells and then iteratively finds the maximum likelihood of cell cluster proportions as well as the cellular response for each gRNA. We demonstrated its performance in a series of simulation experiments. By applying scDecouple to real scCRISPR-seq data, we found that scDecouple enhances the identification of biologically perturbation-related genes. scDecouple can benefit scCRISPR-seq data analysis, especially in the case of heterogeneous samples or complex gRNA libraries. Qiuchen Meng, Lei Wei 0009, Joshua W. K. Ho, Yinqing Li, Xuegong Zhang |
Briefings Bioinform. | 6 |
| 2023 | Translation Rate Prediction and Regulatory Motif Discovery with Multi-task Learning
Weizhong Zheng, John H. C. Fong, Yuk Kei Wan, Athena H. Y. Chu, Yuanhua Huang, Alan S. L. Wong, Joshua W. K. Ho |
RECOMB | 7 |
| 2022 | Evaluation of Experimental Protocols for Shotgun Whole-Genome Metagenomic Discovery of Antibiotic Resistance GenesabstractShotgun metagenomics has enabled the discovery of antibiotic resistance genes (ARGs). Although there have been numerous studies benchmarking the bioinformatics methods for shotgun metagenomic data analysis, there has not yet been a study that systematically evaluates the performance of different experimental protocols on metagenomic species profiling and ARG detection. In this study, we generated 35 whole genome shotgun metagenomic sequencing data sets for five samples (three human stool and two microbial standard) using seven experimental protocols (KAPA or Flex kits at 50ng, 10ng, or 5ng input amounts; XT kit at 1ng input amount). Using this comprehensive resource, we evaluated the seven protocols in terms of robust detection of ARGs and microbial abundance estimation at various sequencing depths. We found that the data generated by the seven protocols are largely similar. The inter-protocol variability is significantly smaller than the variability between samples or sequencing depths. We found that a sequencing depth of more than 30M is suitable for human stool samples. A higher input amount (50ng) is generally favorable for the KAPA and Flex kits. This systematic benchmarking study sheds light on the impact of sequencing depth, experimental protocol, and DNA input amount on ARG detection in human stool samples. Ken Hung-On Yu, Xiunan Fang, Haobin Yao, Bond Ng, Tak Kwan Leung, Chi Ho Lin, Agnes Sze Wah Chan, Wai Keung Leung, Suet Yi Leung, Joshua W. K. Ho |
IEEE ACM Trans. Comput. Biol. Bioinform. | 11 |
| 2021 | FlowGrid enables fast clustering of very large single-cell RNA-seq dataabstractMOTIVATION: Scalable clustering algorithms are needed to analyze millions of cells in single cell RNA-seq (scRNA-seq) data. RESULTS: Here, we present an open source python package called FlowGrid that can integrate into the Scanpy workflow to perform clustering on very large scRNA-seq datasets. FlowGrid implements a fast density-based clustering algorithm originally designed for flow cytometry data analysis. We introduce a new automated parameter tuning procedure, and show that FlowGrid can achieve comparable clustering accuracy as state-of-the-art clustering algorithms but at a substantially reduced run time for very large single cell RNA-seq datasets. For example, FlowGrid can complete a one-hour clustering task for one million cells in about five min. AVAILABILITY AND IMPLEMENTATION: https://github.com/holab-hku/FlowGrid. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Xiunan Fang, Joshua W. K. Ho |
Bioinform. | 2 |
| 2020 | dv-trio: a family-based variant calling pipeline using DeepVariantabstractMOTIVATION: In 2018, Google published an innovative variant caller, DeepVariant, which converts pileups of sequence reads into images and uses a deep neural network to identify single-nucleotide variants and small insertion/deletions from next-generation sequencing data. This approach outperforms existing state-of-the-art tools. However, DeepVariant was designed to call variants within a single sample. In disease sequencing studies, the ability to examine a family trio (father-mother-affected child) provides greater power for disease mutation discovery. RESULTS: To further improve DeepVariant's variant calling accuracy in family-based sequencing studies, we have developed a family-based variant calling pipeline, dv-trio, which incorporates the trio information from the Mendelian genetic model into variant calling based on DeepVariant. AVAILABILITY AND IMPLEMENTATION: dv-trio is available via an open source BSD3 license at GitHub (https://github.com/VCCRI/dv-trio/). CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Eddie Ip, Clinton Hadinata, Joshua W. K. Ho, Eleni Giannoulatou |
Bioinform. | 3 |
| 2020 | PARC: ultrafast and accurate clustering of phenotypic data of millions of single cellsabstractMOTIVATION: New single-cell technologies continue to fuel the explosive growth in the scale of heterogeneous single-cell data. However, existing computational methods are inadequately scalable to large datasets and therefore cannot uncover the complex cellular heterogeneity. RESULTS: We introduce a highly scalable graph-based clustering algorithm PARC-Phenotyping by Accelerated Refined Community-partitioning-for large-scale, high-dimensional single-cell data (>1 million cells). Using large single-cell flow and mass cytometry, RNA-seq and imaging-based biophysical data, we demonstrate that PARC consistently outperforms state-of-the-art clustering algorithms without subsampling of cells, including Phenograph, FlowSOM and Flock, in terms of both speed and ability to robustly detect rare cell populations. For example, PARC can cluster a single-cell dataset of 1.1 million cells within 13 min, compared with >2 h for the next fastest graph-clustering algorithm. Our work presents a scalable algorithm to cope with increasingly large-scale single-cell analysis. AVAILABILITY AND IMPLEMENTATION: https://github.com/ShobiStassen/PARC. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Shobana V. Stassen, Dickson M. D. Siu, Kelvin C. M. Lee, Joshua W. K. Ho, Hayden Kwok-Hay So, Kevin K. Tsia |
Bioinform. | 4 |
| 2017 | Integrative analysis identifies co-dependent gene expression regulation of BRG1 and CHD7 at distal regulatory sites in embryonic stem cellsabstractMOTIVATION: DNA binding proteins such as chromatin remodellers, transcription factors (TFs), histone modifiers and co-factors often bind cooperatively to activate or repress their target genes in a cell type-specific manner. Nonetheless, the precise role of cooperative binding in defining cell-type identity is still largely uncharacterized. RESULTS: Here, we collected and analyzed 214 public datasets representing chromatin immunoprecipitation followed by sequencing (ChIP-Seq) of 104 DNA binding proteins in embryonic stem cell (ESC) lines. We classified their binding sites into those proximal to gene promoters and those in distal regions, and developed a web resource called Proximal And Distal (PAD) clustering to identify their co-localization at these respective regions. Using this extensive dataset, we discovered an extensive co-localization of BRG1 and CHD7 at distal but not proximal regions. The comparison of co-localization sites to those bound by either BRG1 or CHD7 alone showed an enrichment of ESC master TFs binding and active chromatin architecture at co-localization sites. Most notably, our analysis reveals the co-dependency of BRG1 and CHD7 at distal regions on regulating expression of their common target genes in ESC. This work sheds light on cooperative binding of TF binding proteins in regulating gene expression in ESC, and demonstrates the utility of integrative analysis of a manually curated compendium of genome-wide protein binding profiles in our online resource PAD. AVAILABILITY AND IMPLEMENTATION: PAD is freely available at http://pad.victorchang.edu.au/ and its source code is available via an open source GPL 3.0 license at https://github.com/VCCRI/PAD/. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Pengyi Yang, Andrew J. Oldfield, Taiyun Kim, Andrian Yang, Jean Y. H. Yang, Joshua W. K. Ho |
Bioinform. | 6 |
| 2017 | Falco: a quick and flexible single-cell RNA-seq processing framework on the cloudabstractSummary: Single-cell RNA-seq (scRNA-seq) is increasingly used in a range of biomedical studies. Nonetheless, current RNA-seq analysis tools are not specifically designed to efficiently process scRNA-seq data due to their limited scalability. Here we introduce Falco, a cloud-based framework to enable paralellization of existing RNA-seq processing pipelines using big data technologies of Apache Hadoop and Apache Spark for performing massively parallel analysis of large scale transcriptomic data. Using two public scRNA-seq datasets and two popular RNA-seq alignment/feature quantification pipelines, we show that the same processing pipeline runs 2.6-145.4 times faster using Falco than running on a highly optimized standalone computer. Falco also allows users to utilize low-cost spot instances of Amazon Web Services, providing a ∼65% reduction in cost of analysis. Availability and Implementation: Falco is available via a GNU General Public License at https://github.com/VCCRI/Falco/. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Andrian Yang, Michael Troup, Peijie Lin, Joshua W. K. Ho |
Bioinform. | 4 |
| 2016 | XGSA: A statistical method for cross-species gene set analysisabstractMOTIVATION: Gene set analysis is a powerful tool for determining whether an experimentally derived set of genes is statistically significantly enriched for genes in other pre-defined gene sets, such as known pathways, gene ontology terms, or other experimentally derived gene sets. Current gene set analysis methods do not facilitate comparing gene sets across different organisms as they do not explicitly deal with homology mapping between species. There lacks a systematic investigation about the effect of complex gene homology on cross-species gene set analysis. RESULTS: In this study, we show that not accounting for the complex homology structure when comparing gene sets in two species can lead to false positive discoveries, especially when comparing gene sets that have complex gene homology relationships. To overcome this bias, we propose a straightforward statistical approach, called XGSA, that explicitly takes the cross-species homology mapping into consideration when doing gene set analysis. Simulation experiments confirm that XGSA can avoid false positive discoveries, while maintaining good statistical power compared to other ad hoc approaches for cross-species gene set analysis. We further demonstrate the effectiveness of XGSA with two real-life case studies that aim to discover conserved or species-specific molecular pathways involved in social challenge and vertebrate appendage regeneration. AVAILABILITY AND IMPLEMENTATION: The R source code for XGSA is available under a GNU General Public License at http://github.com/VCCRI/XGSA CONTACT: [email protected]. Djordje Djordjevic, Kenro Kusumi, Joshua W. K. Ho |
Bioinform. | 3 |
| 2015 | hiHMM: Bayesian non-parametric joint inference of chromatin state mapsabstractMOTIVATION: Genome-wide mapping of chromatin states is essential for defining regulatory elements and inferring their activities in eukaryotic genomes. A number of hidden Markov model (HMM)-based methods have been developed to infer chromatin state maps from genome-wide histone modification data for an individual genome. To perform a principled comparison of evolutionarily distant epigenomes, we must consider species-specific biases such as differences in genome size, strength of signal enrichment and co-occurrence patterns of histone modifications. RESULTS: Here, we present a new Bayesian non-parametric method called hierarchically linked infinite HMM (hiHMM) to jointly infer chromatin state maps in multiple genomes (different species, cell types and developmental stages) using genome-wide histone modification data. This flexible framework provides a new way to learn a consistent definition of chromatin states across multiple genomes, thus facilitating a direct comparison among them. We demonstrate the utility of this method using synthetic data as well as multiple modENCODE ChIP-seq datasets. CONCLUSION: The hierarchical and Bayesian non-parametric formulation in our approach is an important extension to the current set of methodologies for comparative chromatin landscape analysis. AVAILABILITY AND IMPLEMENTATION: Source codes are available at https://github.com/kasohn/hiHMM. Chromatin data are available at http://encode-x.med.harvard.edu/data_sets/chromatin/. Kyung-Ah Sohn 0001, Joshua W. K. Ho, Djordje Djordjevic, Hyun-hwan Jeong, Peter J. Park, Ju Han Kim |
Bioinform. | 2 |
| 2014 | Verification and validation of bioinformatics software without a gold standard: a case study of BWA and BowtieabstractBACKGROUND: Bioinformatics software quality assurance is essential in genomic medicine. Systematic verification and validation of bioinformatics software is difficult because it is often not possible to obtain a realistic "gold standard" for systematic evaluation. Here we apply a technique that originates from the software testing literature, namely Metamorphic Testing (MT), to systematically test three widely used short-read sequence alignment programs. RESULTS: MT alleviates the problems associated with the lack of gold standard by checking that the results from multiple executions of a program satisfy a set of expected or desirable properties that can be derived from the software specification or user expectations. We tested BWA, Bowtie and Bowtie2 using simulated data and one HapMap dataset. It is interesting to observe that multiple executions of the same aligner using slightly modified input FASTQ sequence file, such as after randomly re-ordering of the reads, may affect alignment results. Furthermore, we found that the list of variant calls can be affected unless strict quality control is applied during variant calling. CONCLUSION: Thorough testing of bioinformatics software is important in delivering clinical genomic medicine. This paper demonstrates a different framework to test a program that involves checking its properties, thus greatly expanding the number and repertoire of test cases we can apply in practice. Eleni Giannoulatou, Shin-Ho Park, David T. Humphreys, Joshua W. K. Ho |
BMC Bioinform. | 4 |
| 2011 | Gene-gene interaction filtering with ensemble of filtersabstractBACKGROUND: Complex diseases are commonly caused by multiple genes and their interactions with each other. Genome-wide association (GWA) studies provide us the opportunity to capture those disease associated genes and gene-gene interactions through panels of SNP markers. However, a proper filtering procedure is critical to reduce the search space prior to the computationally intensive gene-gene interaction identification step. In this study, we show that two commonly used SNP-SNP interaction filtering algorithms, ReliefF and tuned ReliefF (TuRF), are sensitive to the order of the samples in the dataset, giving rise to unstable and suboptimal results. However, we observe that the 'unstable' results from multiple runs of these algorithms can provide valuable information about the dataset. We therefore hypothesize that aggregating results from multiple runs of the algorithm may improve the filtering performance. RESULTS: We propose a simple and effective ensemble approach in which the results from multiple runs of an unstable filter are aggregated based on the general theory of ensemble learning. The ensemble versions of the ReliefF and TuRF algorithms, referred to as ReliefF-E and TuRF-E, are robust to sample order dependency and enable a more informative investigation of data characteristics. Using simulated and real datasets, we demonstrate that both the ensemble of ReliefF and the ensemble of TuRF can generate a much more stable SNP ranking than the original algorithms. Furthermore, the ensemble of TuRF achieved the highest success rate in comparison to many state-of-the-art algorithms as well as traditional χ2-test and odds ratio methods in terms of retaining gene-gene interactions. Pengyi Yang, Joshua W. K. Ho, Jean Y. H. Yang, Bing Bing Zhou |
BMC Bioinform. | 2 |
| 2011 | Testing and validating machine learning classifiers by metamorphic testing
Xiaoyuan Xie, Joshua W. K. Ho, Christian Murphy, Gail E. Kaiser, Baowen Xu, Tsong Yueh Chen |
J. Syst. Softw. | 2 |
| 2010 | A genetic ensemble approach for gene-gene interaction identificationabstractBACKGROUND: It has now become clear that gene-gene interactions and gene-environment interactions are ubiquitous and fundamental mechanisms for the development of complex diseases. Though a considerable effort has been put into developing statistical models and algorithmic strategies for identifying such interactions, the accurate identification of those genetic interactions has been proven to be very challenging. METHODS: In this paper, we propose a new approach for identifying such gene-gene and gene-environment interactions underlying complex diseases. This is a hybrid algorithm and it combines genetic algorithm (GA) and an ensemble of classifiers (called genetic ensemble). Using this approach, the original problem of SNP interaction identification is converted into a data mining problem of combinatorial feature selection. By collecting various single nucleotide polymorphisms (SNP) subsets as well as environmental factors generated in multiple GA runs, patterns of gene-gene and gene-environment interactions can be extracted using a simple combinatorial ranking method. Also considered in this study is the idea of combining identification results obtained from multiple algorithms. A novel formula based on pairwise double fault is designed to quantify the degree of complementarity. CONCLUSIONS: Our simulation study demonstrates that the proposed genetic ensemble algorithm has comparable identification power to Multifactor Dimensionality Reduction (MDR) and is slightly better than Polymorphism Interaction Analysis (PIA), which are the two most popular methods for gene-gene interaction identification. More importantly, the identification results generated by using our genetic ensemble algorithm are highly complementary to those obtained by PIA and MDR. Experimental results from our simulation studies and real world data application also confirm the effectiveness of the proposed genetic ensemble algorithm, as well as the potential benefits of combining identification results from different algorithms. Pengyi Yang, Joshua W. K. Ho, Albert Y. Zomaya, Bing Bing Zhou |
BMC Bioinform. | 2 |
| 2009 | An innovative approach for testing bioinformatics programs using metamorphic testingabstractBACKGROUND: Recent advances in experimental and computational technologies have fueled the development of many sophisticated bioinformatics programs. The correctness of such programs is crucial as incorrectly computed results may lead to wrong biological conclusion or misguided downstream experimentation. Common software testing procedures involve executing the target program with a set of test inputs and then verifying the correctness of the test outputs. However, due to the complexity of many bioinformatics programs, it is often difficult to verify the correctness of the test outputs. Therefore our ability to perform systematic software testing is greatly hindered. RESULTS: We propose to use a novel software testing technique, metamorphic testing (MT), to test a range of bioinformatics programs. Instead of requiring a mechanism to verify whether an individual test output is correct, the MT technique verifies whether a pair of test outputs conform to a set of domain specific properties, called metamorphic relations (MRs), thus greatly increases the number and variety of test cases that can be applied. To demonstrate how MT is used in practice, we applied MT to test two open-source bioinformatics programs, namely GNLab and SeqMap. In particular we show that MT is simple to implement, and is effective in detecting faults in a real-life program and some artificially fault-seeded programs. Further, we discuss how MT can be applied to test programs from various domains of bioinformatics. CONCLUSION: This paper describes the application of a simple, effective and automated technique to systematically test a range of bioinformatics programs. We show how MT can be implemented in practice through two real-life case studies. Since many bioinformatics programs, particularly those for large scale simulation and data analysis, are hard to test systematically, their developers may benefit from using MT as part of the testing strategy. Therefore our work represents a significant step towards software reliability in bioinformatics. Tsong Yueh Chen, Joshua W. K. Ho, Huai Liu, Xiaoyuan Xie |
BMC Bioinform. | 2 |
| 2009 | A voting approach to identify a small number of highly predictive genes using multiple classifiersabstractBACKGROUND: Microarray gene expression profiling has provided extensive datasets that can describe characteristics of cancer patients. An important challenge for this type of data is the discovery of gene sets which can be used as the basis of developing a clinical predictor for cancer. It is desirable that such gene sets be compact, give accurate predictions across many classifiers, be biologically relevant and have good biological process coverage. RESULTS: By using a new type of multiple classifier voting approach, we have identified gene sets that can predict breast cancer prognosis accurately, for a range of classification algorithms. Unlike a wrapper approach, our method is not specialised towards a single classification technique. Experimental analysis demonstrates higher prediction accuracies for our sets of genes compared to previous work in the area. Moreover, our sets of genes are generally more compact than those previously proposed. Taking a biological viewpoint, from the literature, most of the genes in our sets are known to be strongly related to cancer. CONCLUSION: We show that it is possible to obtain superior classification accuracy with our approach and obtain a compact gene set that is also biologically relevant and has good coverage of different biological processes. Md. Rafiul Hassan, M. Maruf Hossain, James Bailey 0001, Geoff MacIntyre, Joshua W. K. Ho, Kotagiri Ramamohanarao |
BMC Bioinform. | 5 |
| 2008 | Differential variability analysis of gene expression and its application to human diseasesabstractMOTIVATION: Current microarray analyses focus on identifying sets of genes that are differentially expressed (DE) or differentially coexpressed (DC) in different biological states (e.g. diseased versus non-diseased). We observed that in many human diseases, some genes have a significant increase or decrease in expression variability (variance). As these observed changes in expression variability may be caused by alteration of the underlying expression dynamics, such differential variability (DV) patterns are also biologically interesting. RESULTS: Here we propose a novel analysis for changes in gene expression variability between groups of samples, which we call differential variability analysis. We introduce the concept of differential variability (DV), and present a simple procedure for identifying DV genes from microarray data. Our procedure is evaluated with simulated and real microarray datasets. The effect of data preprocessing methods on identification of DV gene is investigated. The biological significance of DV analysis is demonstrated with four human disease datasets. The relationships among DV, DE and DC genes are investigated. The results suggest that changes in expression variability are associated with changes in coexpression pattern, which imply that DV is not merely stochastic noise, but informative signal. AVAILABILITY: The R source code for differential variability analysis is available from the contact authors upon request. Joshua W. K. Ho, Maurizio Stefani, Cristobal G. dos Remedios, Michael A. Charleston |
ISMB | 1 |
| 2006 | SeqVis: Visualization of compositional heterogeneity in large alignments of nucleotidesabstractUNLABELLED: Most phylogenetic methods assume that the sequences evolved under homogeneous, stationary and reversible conditions. Compositional heterogeneity in data intended for studies of phylogeny suggests that the data did not evolve under these conditions. SeqVis, a Java application for analysis of nucleotide content, reads sequence alignments in several formats and plots the nucleotide content in a tetrahedron. Once plotted, outliers can be identified, thus allowing for decisions on the applicability of the data for phylogenetic analysis. AVAILABILITY: http://www.bio.usyd.edu.au/jermiin/programs.htm. Joshua W. K. Ho, Cameron E. Adams, Jie Bin Lew, Timothy J. Matthews, Chiu Chin Ng, Arash Shahabi-Sirjani, Leng Hong Tan, Simon Easteal, Susan R. Wilson, Lars S. Jermiin |
Bioinform. | 1 |
| 2005 | GEOMI: GEOmetry for Maximum Insight
Adel Ahmed, Tim Dwyer, Michael Forster, Xiaoyan Fu, Joshua W. K. Ho, Seok-Hee Hong 0001, Dirk Koschützki, Colin Murray, Nikola S. Nikolov, Ronnie Taib, Alexandre Tarassov, Kai Xu 0003 |
GD | 5 |
| 2005 | Drawing Clustered Graphs in Three Dimensions
Joshua W. K. Ho, Seok-Hee Hong 0001 |
GD | 1 |