Desmond G. Higgins

dblp:89/6822 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
0since 2021 · last 2020
0000-0002-3952-3285ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 28 · 5 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
22 papers
Bioinformatics and computational biology · 98% Computational science and engineering · 1% Environmental and earth informatics · 1%

Topics — the 28 heaviest of 28, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
multiple sequence alignment
1.092020
QuanTest2: benchmarking multiple sequence alignments using secondary structure prediction · Bioinform. 2020
Using de novo protein structure predictions to measure the quality of very large multiple sequence alignments · Bioinform. 2016
Making automated multiple alignments of very large numbers of protein sequences · Bioinform. 2013
Bioinformatics and computational biology › multiple sequence alignment
alignment quality assessment
0.942020
QuanTest2: benchmarking multiple sequence alignments using secondary structure prediction · Bioinform. 2020
Protein multiple sequence alignment benchmarking through secondary structure prediction · Bioinform. 2017
Making automated multiple alignments of very large numbers of protein sequences · Bioinform. 2013
Bioinformatics and computational biology
sequence analysis
0.7102020
QuanTest2: benchmarking multiple sequence alignments using secondary structure prediction · Bioinform. 2020
Making automated multiple alignments of very large numbers of protein sequences · Bioinform. 2013
Clustal W and Clustal X version 2.0 · Bioinform. 2007
Bioinformatics and computational biology
protein structure prediction
0.322017
Using de novo protein structure predictions to measure the quality of very large multiple sequence alignments · Bioinform. 2016
Protein multiple sequence alignment benchmarking through secondary structure prediction · Bioinform. 2017
Bioinformatics and computational biology › multiple sequence alignment
protein multiple sequence alignment
0.312017
Protein multiple sequence alignment benchmarking through secondary structure prediction · Bioinform. 2017
Bioinformatics and computational biology › multiple sequence alignment › alignment quality assessment
alignment benchmark
0.212016
Using de novo protein structure predictions to measure the quality of very large multiple sequence alignments · Bioinform. 2016
Bioinformatics and computational biology › gene expression analysis
microarray data analysis
0.122005
MADE4: an R package for multivariate analysis of gene expression data · Bioinform. 2005
Between-group analysis of microarray data · Bioinform. 2002
Bioinformatics and computational biology › protein structure prediction
secondary structure prediction
0.112017
Protein multiple sequence alignment benchmarking through secondary structure prediction · Bioinform. 2017
Bioinformatics and computational biology
sequence alignment
0.141999
MIAH: automatic alignment of eukaryotic SSU rRNAs · Bioinform. 1999
Optimization of ribosomal RNA profile alignments · Bioinform. 1998
Empirical estimation of the reliability of ribosomal RNA alignments · Bioinform. 1998
Bioinformatics and computational biology
gene expression analysis
0.112007
Integrating transcription factor binding site information with gene expression datasets · Bioinform. 2007
Bioinformatics and computational biology › sequence analysis
motif discovery
0.112007
Integrating transcription factor binding site information with gene expression datasets · Bioinform. 2007
Bioinformatics and computational biology
protein structure analysis
0.122006
APDB: a web server to evaluate the accuracy of sequence alignments using structural information · Bioinform. 2006
OBSTRUCT: a program to obtain largest cliques from a protein sequence set according to structural resolution and sequence similarity · Comput. Appl. Biosci. 1992
Bioinformatics and computational biology › sequence alignment
sequence alignment evaluation
0.112006
APDB: a web server to evaluate the accuracy of sequence alignments using structural information · Bioinform. 2006
Bioinformatics and computational biology › multiple sequence alignment
progressive alignment
0.122005
Evaluation of iterative alignment algorithms for multiple alignment · Bioinform. 2005
Fast and sensitive multiple sequence alignments on a microcomputer · Comput. Appl. Biosci. 1989
Bioinformatics and computational biology › multiple sequence alignment
iterative alignment
0.112005
Evaluation of iterative alignment algorithms for multiple alignment · Bioinform. 2005
Computational science and engineering › numerical linear algebra
iterative refinement
0.112005
Evaluation of iterative alignment algorithms for multiple alignment · Bioinform. 2005
Bioinformatics and computational biology › structural bioinformatics
protein structure
0.012013
Making automated multiple alignments of very large numbers of protein sequences · Bioinform. 2013
Environmental and earth informatics
ordination
0.012002
Between-group analysis of microarray data · Bioinform. 2002
Bioinformatics and computational biology
structural bioinformatics
0.022006
APDB: a web server to evaluate the accuracy of sequence alignments using structural information · Bioinform. 2006
OBSTRUCT: a program to obtain largest cliques from a protein sequence set according to structural resolution and sequence similarity · Comput. Appl. Biosci. 1992
Bioinformatics and computational biology › sequence analysis › sequence similarity search
sequence database search
0.021994
Improved sensitivity of profile searches through the use of sequence weights and gap excision · Comput. Appl. Biosci. 1994
EMBLSCAN: fast approximate DNA database searches on compact disc · Comput. Appl. Biosci. 1992
Bioinformatics and computational biology › sequence analysis › sequence profile analysis
profile search
0.011994
Improved sensitivity of profile searches through the use of sequence weights and gap excision · Comput. Appl. Biosci. 1994
Bioinformatics and computational biology › genome annotation
gene prediction
0.011992
GCWIND: a microcomputer program for identifying open reading frames according to codon positional G+C content · Comput. Appl. Biosci. 1992
Bioinformatics and computational biology › genome annotation › coding region prediction
open reading frame identification
0.011992
GCWIND: a microcomputer program for identifying open reading frames according to codon positional G+C content · Comput. Appl. Biosci. 1992
Bioinformatics and computational biology
phylogenetics
0.031992
CLUSTAL V: improved software for multiple sequence alignment · Comput. Appl. Biosci. 1992
Sequence ordinations: a multivariate analysis approach to analysing large sequence data sets · Comput. Appl. Biosci. 1992
Fast and sensitive multiple sequence alignments on a microcomputer · Comput. Appl. Biosci. 1989
Bioinformatics and computational biology › sequence analysis
sequence weighting
0.011998
Optimization of ribosomal RNA profile alignments · Bioinform. 1998
Bioinformatics and computational biology › phylogenetics
phylogenetic inference
0.021992
CLUSTAL V: improved software for multiple sequence alignment · Comput. Appl. Biosci. 1992
Fast and sensitive multiple sequence alignments on a microcomputer · Comput. Appl. Biosci. 1989
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.011987
Interfacing similarity search software with the sequence retrieval system ACNUC · Comput. Appl. Biosci. 1987
Bioinformatics and computational biology › sequence analysis › database search
sequence retrieval
0.011987
Interfacing similarity search software with the sequence retrieval system ACNUC · Comput. Appl. Biosci. 1987

Methods — techniques the papers use, named apart from their topics

sum-of-pairs score · 0.4benchmark comparison · 0.3de novo protein structure prediction · 0.2chained guide trees · 0.2iterative refinement · 0.2guide tree construction · 0.2progressive alignment · 0.1correspondence analysis · 0.1between-group analysis · 0.1co-inertia analysis · 0.1
YearPublicationVenuePosition
2020 QuanTest2: benchmarking multiple sequence alignments using secondary structure prediction
abstract
MOTIVATION: Secondary structure prediction accuracy (SSPA) in the QuanTest benchmark can be used to measure accuracy of a multiple sequence alignment. SSPA correlates well with the sum-of-pairs score, if the results are averaged over many alignments but not on an alignment-by-alignment basis. This is due to a sub-optimal selection of reference and non-reference sequences in QuanTest. RESULTS: We develop an improved strategy for selecting reference and non-reference sequences for a new benchmark, QuanTest2. In QuanTest2, SSPA and SP correlate better on an alignment-by-alignment basis than in QuanTest. Guide-trees for QuanTest2 are more balanced with respect to reference sequences than in QuanTest. QuanTest2 scores correlate well with other well-established benchmarks. AVAILABILITY AND IMPLEMENTATION: QuanTest2 is available at http://bioinf.ucd.ie/quantest2.tar, comprises of reference and non-reference sequence sets and a scoring script. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fabian Sievers, Desmond G. Higgins
Bioinform.2
2017 Protein multiple sequence alignment benchmarking through secondary structure prediction
abstract
Motivation: Multiple sequence alignment (MSA) is commonly used to analyze sets of homologous protein or DNA sequences. This has lead to the development of many methods and packages for MSA over the past 30 years. Being able to compare different methods has been problematic and has relied on gold standard benchmark datasets of 'true' alignments or on MSA simulations. A number of protein benchmark datasets have been produced which rely on a combination of manual alignment and/or automated superposition of protein structures. These are either restricted to very small MSAs with few sequences or require manual alignment which can be subjective. In both cases, it remains very difficult to properly test MSAs of more than a few dozen sequences. PREFAB and HomFam both rely on using a small subset of sequences of known structure and do not fairly test the quality of a full MSA. Results: In this paper we describe QuanTest, a fully automated and highly scalable test system for protein MSAs which is based on using secondary structure prediction accuracy (SSPA) to measure alignment quality. This is based on the assumption that better MSAs will give more accurate secondary structure predictions when we include sequences of known structure. SSPA measures the quality of an entire alignment however, not just the accuracy on a handful of selected sequences. It can be scaled to alignments of any size but here we demonstrate its use on alignments of either 200 or 1000 sequences. This allows the testing of slow accurate programs as well as faster, less accurate ones. We show that the scores from QuanTest are highly correlated with existing benchmark scores. We also validate the method by comparing a wide range of MSA alignment options and by including different levels of mis-alignment into MSA, and examining the effects on the scores. Availability and Implementation: QuanTest is available from http://www.bioinf.ucd.ie/download/QuanTest.tgz. Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online.
Quan Le, Fabian Sievers, Desmond G. Higgins
Bioinform.3
2016 Using de novo protein structure predictions to measure the quality of very large multiple sequence alignments
abstract
MOTIVATION: Multiple sequence alignments (MSAs) with large numbers of sequences are now commonplace. However, current multiple alignment benchmarks are ill-suited for testing these types of alignments, as test cases either contain a very small number of sequences or are based purely on simulation rather than empirical data. RESULTS: We take advantage of recent developments in protein structure prediction methods to create a benchmark (ContTest) for protein MSAs containing many thousands of sequences in each test case and which is based on empirical biological data. We rank popular MSA methods using this benchmark and verify a recent result showing that chained guide trees increase the accuracy of progressive alignment packages on datasets with thousands of proteins. AVAILABILITY AND IMPLEMENTATION: Benchmark data and scripts are available for download at http://www.bioinf.ucd.ie/download/ContTest.tar.gz CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Gearoid Fox, Fabian Sievers, Desmond G. Higgins
Bioinform.3
2015 OD-seq: outlier detection in multiple sequence alignments
abstract
BACKGROUND: Multiple sequence alignments (MSA) are widely used in sequence analysis for a variety of tasks. Outlier sequences can make downstream analyses unreliable or make the alignments less accurate while they are being constructed. This paper describes a simple method for automatically detecting outliers and accompanying software called OD-seq. It is based on finding sequences whose average distance to the rest of the sequences in a dataset, is anomalous. RESULTS: The software can take a MSA, distance matrix or set of unaligned sequences as input. Outlier sequences are found by examining the average distance of each sequence to the rest. Anomalous average distances are then found using the interquartile range of the distribution of average distances or by bootstrapping them. The complexity of any analysis of a distance matrix is normally at least O(N(2)) for N sequences. This is prohibitive for large N but is reduced here by using the mBed algorithm from Clustal Omega. This reduces the complexity to O(N log(N)) which makes even very large alignments easy to analyse on a single core. We tested the ability of OD-seq to detect outliers using artificial test cases of sequences from Pfam families, seeded with sequences from other Pfam families. Using a MSA as input, OD-seq is able to detect outliers with very high sensitivity and specificity. CONCLUSION: OD-seq is a practical and simple method to detect outliers in MSAs. It can also detect outliers in sets of unaligned sequences, but with reduced accuracy. For medium sized alignments, of a few thousand sequences, it can detect outliers in a few seconds. Software available as http://www.bioinf.ucd.ie/download/od-seq.tar.gz.
Peter Jehl, Fabian Sievers, Desmond G. Higgins
BMC Bioinform.3
2014 Systematic exploration of guide-tree topology effects for small protein alignments
abstract
BACKGROUND: Guide-trees are used as part of an essential heuristic to enable the calculation of multiple sequence alignments. They have been the focus of much method development but there has been little effort at determining systematically, which guide-trees, if any, give the best alignments. Some guide-tree construction schemes are based on pair-wise distances amongst unaligned sequences. Others try to emulate an underlying evolutionary tree and involve various iteration methods. RESULTS: We explore all possible guide-trees for a set of protein alignments of up to eight sequences. We find that pairwise distance based default guide-trees sometimes outperform evolutionary guide-trees, as measured by structure derived reference alignments. However, default guide-trees fall way short of the optimum attainable scores. On average chained guide-trees perform better than balanced ones but are not better than default guide-trees for small alignments. CONCLUSIONS: Alignment methods that use Consistency or hidden Markov models to make alignments are less susceptible to sub-optimal guide-trees than simpler methods, that basically use conventional sequence alignment between profiles. The latter appear to be affected positively by evolutionary based guide-trees for difficult alignments and negatively for easy alignments. One phylogeny aware alignment program can strongly discriminate between good and bad guide-trees. The results for randomly chained guide-trees improve with the number of sequences.
Fabian Sievers, Graham M. Hughes, Desmond G. Higgins
BMC Bioinform.3
2013 Making automated multiple alignments of very large numbers of protein sequences
abstract
MOTIVATION: Recent developments in sequence alignment software have made possible multiple sequence alignments (MSAs) of >100 000 sequences in reasonable times. At present, there are no systematic analyses concerning the scalability of the alignment quality as the number of aligned sequences is increased. RESULTS: We benchmarked a wide range of widely used MSA packages using a selection of protein families with some known structures and found that the accuracy of such alignments decreases markedly as the number of sequences grows. This is more or less true of all packages and protein families. The phenomenon is mostly due to the accumulation of alignment errors, rather than problems in guide-tree construction. This is partly alleviated by using iterative refinement or selectively adding sequences. The average accuracy of progressive methods by comparison with structure-based benchmarks can be improved by incorporating information derived from high-quality structural alignments of sequences with solved structures. This suggests that the availability of high quality curated alignments will have to complement algorithmic and/or software developments in the long-term. AVAILABILITY AND IMPLEMENTATION: Benchmark data used in this study are available at http://www.clustal.org/omega/homfam-20110613-25.tar.gz and http://www.clustal.org/omega/bali3fam-26.tar.gz. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Fabian Sievers, David Dineen, Andreas Wilm, Desmond G. Higgins
Bioinform.4
2010 Detecting microRNA activity from gene expression data
abstract
BACKGROUND: MicroRNAs (miRNAs) are non-coding RNAs that regulate gene expression by binding to the messenger RNA (mRNA) of protein coding genes. They control gene expression by either inhibiting translation or inducing mRNA degradation. A number of computational techniques have been developed to identify the targets of miRNAs. In this study we used predicted miRNA-gene interactions to analyse mRNA gene expression microarray data to predict miRNAs associated with particular diseases or conditions. RESULTS: Here we combine correspondence analysis, between group analysis and co-inertia analysis (CIA) to determine which miRNAs are associated with differences in gene expression levels in microarray data sets. Using a database of miRNA target predictions from TargetScan, TargetScanS, PicTar4way PicTar5way, and miRanda and combining these data with gene expression levels from sets of microarrays, this method produces a ranked list of miRNAs associated with a specified split in samples. We applied this to three different microarray datasets, a papillary thyroid carcinoma dataset, an in-house dataset of lipopolysaccharide treated mouse macrophages, and a multi-tissue dataset. In each case we were able to identified miRNAs of biological importance. CONCLUSIONS: We describe a technique to integrate gene expression data and miRNA target predictions from multiple sources.
Stephen F. Madden, Susan B. Carpenter, Ian B. Jeffery, Harry Björkbacka, Katherine A. Fitzgerald, Luke A. O'Neill, Desmond G. Higgins
BMC Bioinform.7
2007 Integrating transcription factor binding site information with gene expression datasets
abstract
MOTIVATION: Microarrays are widely used to measure gene expression differences between sets of biological samples. Many of these differences will be due to differences in the activities of transcription factors. In principle, these differences can be detected by associating motifs in promoters with differences in gene expression levels between the groups. In practice, this is hard to do. RESULTS: We combine correspondence analysis, between group analysis and co-inertia analysis to determine which motifs, from a database of promoter motifs, are strongly associated with differences in gene expression levels. Given a database of motifs and gene expression levels from a set of arrays, the method produces a ranked list of motifs associated with any specified split in the arrays. We give an example using the Gene Atlas compendium of gene expression levels for human tissues where we search for motifs that are associated with expression in central nervous system (CNS) or muscle tissues. Most of the motifs that we find are known from previous work to be strongly associated with expression in CNS or muscle. We give a second example using a published prostate cancer dataset where we can simply and clearly find which transcriptional pathways are associated with differences between benign and metastatic samples. AVAILABILITY: The source code is freely available upon request from the authors.
Ian B. Jeffery, Stephen F. Madden, Paul A. McGettigan, Guy Perrière, Aedín C. Culhane, Desmond G. Higgins
Bioinform.6
2007 Clustal W and Clustal X version 2.0
abstract
SUMMARY: The Clustal W and Clustal X multiple sequence alignment programs have been completely rewritten in C++. This will facilitate the further development of the alignment algorithms in the future and has allowed proper porting of the programs to the latest versions of Linux, Macintosh and Windows operating systems. AVAILABILITY: The programs can be run on-line from the EBI web server: http://www.ebi.ac.uk/tools/clustalw2. The source code and executables for Windows, Linux and Macintosh computers are available from the EBI ftp site ftp://ftp.ebi.ac.uk/pub/software/clustalw2/
Mark A. Larkin, Gordon Blackshields, Nigel P. Brown, R. Chenna, Paul A. McGettigan, Hamish McWilliam, Franck Valentin, Iain M. Wallace, Andreas Wilm, Rodrigo Lopez, Julie Dawn Thompson, Toby J. Gibson, Desmond G. Higgins
Bioinform.13
2007 Supervised multivariate analysis of sequence groups to identify specificity determining residues
abstract
BACKGROUND: Proteins that evolve from a common ancestor can change functionality over time, and it is important to be able identify residues that cause this change. In this paper we show how a supervised multivariate statistical method, Between Group Analysis (BGA), can be used to identify these residues from families of proteins with different substrate specifities using multiple sequence alignments. RESULTS: We demonstrate the usefulness of this method on three different test cases. Two of these test cases, the Lactate/Malate dehydrogenase family and Nucleotidyl Cyclases, consist of two functional groups. The other family, Serine Proteases consists of three groups. BGA was used to analyse and visualise these three families using two different encoding schemes for the amino acids. CONCLUSION: This overall combination of methods in this paper is powerful and flexible while being computationally very fast and simple. BGA is especially useful because it can be used to analyse any number of functional classes. In the examples we used in this paper, we have only used 2 or 3 classes for demonstration purposes but any number can be used and visualised.
Iain M. Wallace, Desmond G. Higgins
BMC Bioinform.2
2006 APDB: a web server to evaluate the accuracy of sequence alignments using structural information
abstract
UNLABELLED: The APDB webserver uses structural information to evaluate the alignment of sequences with known structures. It returns a score correlated to the overall alignment accuracy as well as a local evaluation. Any sequence alignment can be analyzed with APDB provided it includes at least two proteins with known structures. Sequences without a known structure are simply ignored and do not contribute to the scoring procedure. AVAILABILITY: APDB is part of the T-Coffee suite of tools for alignment analysis, it is available on www.tcoffee.org. A stand-alone version of the package is also available as a freeware open source from the same address.
Fabrice Armougom, Olivier Poirot, Sébastien Moretti, Desmond G. Higgins, Phillip Bucher, Vladimir Keduas, Cédric Notredame
Bioinform.4
2006 Comparison and evaluation of methods for generating differentially expressed gene lists from microarray data
abstract
BACKGROUND: Numerous feature selection methods have been applied to the identification of differentially expressed genes in microarray data. These include simple fold change, classical t-statistic and moderated t-statistics. Even though these methods return gene lists that are often dissimilar, few direct comparisons of these exist. We present an empirical study in which we compare some of the most commonly used feature selection methods. We apply these to 9 publicly available datasets, and compare, both the gene lists produced and how these perform in class prediction of test datasets. RESULTS: In this study, we compared the efficiency of the feature selection methods; significance analysis of microarrays (SAM), analysis of variance (ANOVA), empirical bayes t-statistic, template matching, maxT, between group analysis (BGA), Area under the receiver operating characteristic (ROC) curve, the Welch t-statistic, fold change, rank products, and sets of randomly selected genes. In each case these methods were applied to 9 different binary (two class) microarray datasets. Firstly we found little agreement in gene lists produced by the different methods. Only 8 to 21% of genes were in common across all 10 feature selection methods. Secondly, we evaluated the class prediction efficiency of each gene list in training and test cross-validation using four supervised classifiers. CONCLUSION: We report that the choice of feature selection method, the number of genes in the genelist, the number of cases (samples) and the noise in the dataset, substantially influence classification success. Recommendations are made for choice of feature selection. Area under a ROC curve performed well with datasets that had low levels of noise and large sample size. Rank products performs well when datasets had low numbers of samples or high levels of noise. The Empirical bayes t-statistic performed well across a range of sample sizes.
Ian B. Jeffery, Desmond G. Higgins, Aedín C. Culhane
BMC Bioinform.2
2005 MADE4: an R package for multivariate analysis of gene expression data
abstract
Summary: MADE4, microarray ade4, is a software package that facilitates multivariate analysis of microarray gene-expression data. MADE4 accepts a wide variety of gene-expression data formats. MADE4 takes advantage of the extensive multivariate statistical and graphical functions in the R package ade4, extending these for application to microarray data. In addition, MADE4 provides new graphical and visualization tools that aid in interpretation of multivariate analysis of microarray data. Availability: The R package MADE4 is available from Bioconductor http://bioinf.vcd.ie/software and from Bioconductor http://www.bioconductor.org Contact: [email protected] Supplementary information: MADE4 is well documented. There are tutorials, in the form of vignettes, which describe typical analyses. In addition, the MADE4 manual provides descriptions and examples for each function.
Aedín C. Culhane, Jean Thioulouse, Guy Perrière, Desmond G. Higgins
Bioinform.4
2005 Evaluation of iterative alignment algorithms for multiple alignment
abstract
MOTIVATION: Iteration has been used a number of times as an optimization method to produce multiple alignments, either alone or in combination with other methods. Iteration has a great advantage in that it is often very simple both in terms of coding the algorithms and the complexity of the time and memory requirements. In this paper, we systematically test several different iteration strategies by comparing the results on sets of alignment test cases. RESULTS: We tested three schemes where iteration is used to improve an existing alignment. This was found to be remarkably effective and could induce a significant improvement in the accuracy of alignments from most packages. For example the average accuracy of ClustalW was improved by over 6% on the hardest test cases. Iteration was found to be even more powerful when it was directly incorporated into a progressive alignment scheme. Here, iteration was used to improve subalignments at each step of progressive alignment. The beneficial effects of iteration come, in part, from the ability to get round the usual local minimum problem with progressive alignment. This ability can also be used to help reduce the complexity of T-Coffee, without losing accuracy. Alignments can be generated, using T-Coffee, to align subgroups of sequences, which can then be iteratively improved and merged. AVAILABILITY: All of the scripts are freely available on the web at http://www.bioinf.ucd.ie/people/iain/iteration.html CONTACT: [email protected].
Iain M. Wallace, Orla O'Sullivan, Desmond G. Higgins
Bioinform.3
2003 A SAT-Based Approach to Multiple Sequence Alignment
Steven D. Prestwich, Desmond G. Higgins, Orla O'Sullivan
CP2
2003 Cross-platform comparison and visualisation of gene expression data using co-inertia analysis
abstract
BACKGROUND: Rapid development of DNA microarray technology has resulted in different laboratories adopting numerous different protocols and technological platforms, which has severely impacted on the comparability of array data. Current cross-platform comparison of microarray gene expression data are usually based on cross-referencing the annotation of each gene transcript represented on the arrays, extracting a list of genes common to all arrays and comparing expression data of this gene subset. Unfortunately, filtering of genes to a subset represented across all arrays often excludes many thousands of genes, because different subsets of genes from the genome are represented on different arrays. We wish to describe the application of a powerful yet simple method for cross-platform comparison of gene expression data. Co-inertia analysis (CIA) is a multivariate method that identifies trends or co-relationships in multiple datasets which contain the same samples. CIA simultaneously finds ordinations (dimension reduction diagrams) from the datasets that are most similar. It does this by finding successive axes from the two datasets with maximum covariance. CIA can be applied to datasets where the number of variables (genes) far exceeds the number of samples (arrays) such is the case with microarray analyses. RESULTS: We illustrate the power of CIA for cross-platform analysis of gene expression data by using it to identify the main common relationships in expression profiles on a panel of 60 tumour cell lines from the National Cancer Institute (NCI) which have been subjected to microarray studies using both Affymetrix and spotted cDNA array technology. The co-ordinates of the CIA projections of the cell lines from each dataset are graphed in a bi-plot and are connected by a line, the length of which indicates the divergence between the two datasets. Thus, CIA provides graphical representation of consensus and divergence between the gene expression profiles from different microarray platforms. Secondly, the genes that define the main trends in the analysis can be easily identified. CONCLUSIONS: CIA is a robust, efficient approach to coupling of gene expression datasets. CIA provides simple graphical representations of the results making it a particularly attractive method for the identification of relationships between large datasets.
Aedín C. Culhane, Guy Perrière, Desmond G. Higgins
BMC Bioinform.3
2002 Between-group analysis of microarray data
abstract
Abstract Motivation: Most supervised classification methods are limited by the requirement for more cases than variables. In microarray data the number of variables (genes) far exceeds the number of cases (arrays), and thus filtering and pre-selection of genes is required. We describe the application of Between Group Analysis (BGA) to the analysis of microarray data. A feature of BGA is that it can be used when the number of variables (genes) exceeds the number of cases (arrays). BGA is based on carrying out an ordination of groups of samples, using a standard method such as Correspondence Analysis (COA), rather than an ordination of the individual microarray samples. As such, it can be viewed as a method of carrying out COA with grouped data. Results: We illustrate the power of the method using two cancer data sets. In both cases, we can quickly and accurately classify test samples from any number of specified a priori groups and identify the genes which characterize these groups. We obtained very high rates of correct classification, as determined by jack-knife or validation experiments with training and test sets. The results are comparable to those from other methods in terms of accuracy but the power and flexibility of BGA make it an especially attractive method for the analysis of microarray cancer data. Availability: The methods described are implemented in ADE-4 which runs under MacOS and Windows, and is freely available at http://pbil.univ-lyon1.fr/ADE-4/. All scripts are available on request. Contact: [email protected] Supplementary information: Supplementary figures and tables are available at http://bioinfo.ucc.ie/BGA/. * To whom correspondence should be addressed.
Aedín C. Culhane, Guy Perrière, Elizabeth C. Considine, Thomas G. Cotter, Desmond G. Higgins
Bioinform.5
1999 MIAH: automatic alignment of eukaryotic SSU rRNAs
abstract
SUMMARY: MIAH is a WWW server for the automatic alignment of new eukaryotic SSU rRNA sequences to an existing alignment of 1500 sequences. AVAILABILITY: http://chah.ucc.ie/MIAH Contact :
Patricia Thébault, Pierre Monestie, Annette McGrath, Desmond G. Higgins
Bioinform.4
1998 COFFEE: an objective function for multiple sequence alignments
abstract
MOTIVATION: In order to increase the accuracy of multiple sequence alignments, we designed a new strategy for optimizing multiple sequence alignments by genetic algorithm. We named it COFFEE (Consistency based Objective Function For alignmEnt Evaluation). The COFFEE score reflects the level of consistency between a multiple sequence alignment and a library containing pairwise alignments of the same sequences. RESULTS: We show that multiple sequence alignments can be optimized for their COFFEE score with the genetic algorithm package SAGA. The COFFEE function is tested on 11 test cases made of structural alignments extracted from 3D_ali. These alignments are compared to those produced using five alternative methods. Results indicate that COFFEE outperforms the other methods when the level of identity between the sequences is low. Accuracy is evaluated by comparison with the structural alignments used as references. We also show that the COFFEE score can be used as a reliability index on multiple sequence alignments. Finally, we show that given a library of structure-based pairwise sequence alignments extracted from FSSP, SAGA can produce high-quality multiple sequence alignments. The main advantage of COFFEE is its flexibility. With COFFEE, any method suitable for making pairwise alignments can be extended to making multiple alignments. AVAILABILITY: The package is available along with the test cases through the WWW: http://www. ebi.ac.uk/cedric CONTACT: [email protected]
Cédric Notredame, Liisa Holm, Desmond G. Higgins
Bioinform.3
1998 Empirical estimation of the reliability of ribosomal RNA alignments
abstract
MOTIVATION: The automatic alignment of rRNA sequences can reproduce manual expert alignments with high, but not perfect, fidelity. We examine the use of empirical methods for the identification of regions of an alignment of a new sequence with an existing large alignment which can confidently be predicted to be correctly aligned. RESULTS: We show how to use a simple jack-knife procedure to derive an estimate of the reliability that is to be expected at each position of a large alignment of eukaryotic rRNA sequences. These reliabilities are then improved using measures that are specific to the input sequence. Regions where the sequence-specific reliability method performs particularly well are identified and seen to correspond with elements in the structure of the rRNA molecules that vary between species in the alignment. We also compare these reliability measures to an algorithmic alignment stability measure. AVAILABILITY: The software is available free of charge by sending an e-mail message to [email protected]. CONTACT: [email protected]
Emmet A. O'Brien, Desmond G. Higgins
Bioinform.2
1998 Optimization of ribosomal RNA profile alignments
abstract
MOTIVATION: Large alignments of ribosomal RNA sequences are maintained at various sites. New sequences are added to these alignments using a combination of manual and automatic methods. We examine the use of profile alignment methods for rRNA alignment and try to optimize the choice of parameters and sequence weights. RESULTS: Using a large alignment of eukaryotic SSU rRNA sequences as a test case, we empirically compared the performance of various sequence weighting schemes over a range of gap penalties. We developed a new weighting scheme which gives most weight to the sequences in the profile that are most similar to the new sequence. We show that it gives the most accurate alignments when combined with a more traditional sequence weighting scheme. AVAILABILITY: The source code of all software is freely available by anonymous ftp from chah.ucc.ie in the directory /home/ftp/pub/emmet,in the compressed file PRNAA.tar: CONTACT: [email protected], [email protected]
Emmet A. O'Brien, Cédric Notredame, Desmond G. Higgins
Bioinform.3
1994 Improved sensitivity of profile searches through the use of sequence weights and gap excision
abstract
Position-specific substitution matrices, known as profiles, derived from multiple sequence alignments are currently used to search sequence databases for distantly related members of protein families. The performance of the database searches is enhanced by using (i) a sequence weighting scheme which assigns higher weights to more distantly related sequences based on branch lengths derived from phylogenetic trees, (ii) exclusion of positions with mainly padding characters at sites of insertions or deletions and (iii) the BLOSUM62 residue comparison matrix. A natural consequence of these modifications is an improvement in the alignment of new sequences to the profiles. However, the accuracy of the alignments can be further increased by employing a similarity residue comparison matrix. These developments are implemented in a program called PROFILEWEIGHT which runs on Unix and Vax computers. The only input required by the program is the multiple sequence alignment. The output from PROFILEWEIGHT is a profile designed to be used by existing searching and alignment programs. Test results from database searches with four different families of proteins show the improved sensitivity of the weighted profiles.
Julie Dawn Thompson, Desmond G. Higgins, Toby J. Gibson
Comput. Appl. Biosci.2
1992 OBSTRUCT: a program to obtain largest cliques from a protein sequence set according to structural resolution and sequence similarity
abstract
A program OBSTRUCT has been developed to obtain the largest possible subset according to specific constraints from a set of protein sequences whose tertiary structures have been determined crystallographically. The user can request a range in sequence similarity level and/or structural resolution. The program optionally includes sequences with known three-dimensional folds elicited from NMR data.
Jaap Heringa, Hubert Sommerfeldt, Desmond G. Higgins, Patrick Argos
Comput. Appl. Biosci.3
1992 Sequence ordinations: a multivariate analysis approach to analysing large sequence data sets
abstract
Ordination is a powerful method for analysing complex data sets but has been largely ignored in sequence analysis. This paper shows how to use principal coordinates analysis to find low-dimensional representations of distance matrices derived from aligned sets of sequences. The method takes a matrix of Euclidean distances between all pairs of sequence and finds a coordinate space where the distances are exactly preserved. The main problem is to find a measure of distance between aligned sequences that is Euclidean. The simplest distance function is the square root of the percentage difference (as measured by identities) between two sequences, where one ignores any positions in the alignment where there is a gap in any sequence. If one does not ignore positions with a gap, the distances cannot be guaranteed to be Euclidean but the deleterious effects are trivial. Two examples of using the method are shown. A set of 226 aligned globins were analysed and the resulting ordination very successfully represents the known patterns of relationship between the sequences. In the other example, a set of 610 aligned 5S rRNA sequences were analysed. Sequence ordinations complement phylogenetic analyses. They should not be viewed as a complete alternative.
Desmond G. Higgins
Comput. Appl. Biosci.1
1992 CLUSTAL V: improved software for multiple sequence alignment
abstract
The CLUSTAL package of multiple sequence alignment programs has been completely rewritten and many new features added. The new software is a single program called CLUSTAL V, which is written in C and can be used on any machine with a standard C compiler. The main new features are the ability to store and reuse old alignments and the ability to calculate phylogenetic trees after alignment. The program is simple to use, completely menu driven and on-line help is provided.
Desmond G. Higgins, Alan J. Bleasby, Rainer Fuchs
Comput. Appl. Biosci.1
1992 EMBLSCAN: fast approximate DNA database searches on compact disc
abstract
An algorithm that allows rapid searching of nucleic acid sequences based on pregenerated index files is described. The programs and index files for searching the entire EMBL nucleotide sequence collection are being distributed on the EMBL Data Library's CD-ROM.
Desmond G. Higgins, Peter Stoehr
Comput. Appl. Biosci.1
1992 GCWIND: a microcomputer program for identifying open reading frames according to codon positional G+C content
abstract
GCWIND is a microcomputer (IBM-PC compatible) program for the identification of protein-coding open reading frames. The program is similar to the FRAME program, but the latter has only been implemented for a specialized graphics package. The base compositions (%G+C) for each of the three possible reading phases through the DNA sequence are displayed separately, together with the positions of potential translation initiation and termination codons (on the leading and complementary strands), to provide an immediate representation of those regions within the sequence that have coding potential.
Denis C. Shields, Desmond G. Higgins, P. M. Sharp
Comput. Appl. Biosci.2
1989 Fast and sensitive multiple sequence alignments on a microcomputer
abstract
A strategy is described for the rapid alignment of many long nucleic acid or protein sequences on a microcomputer. The program described can handle up to 100 sequences of 1200 residues each. The approach is based on progressively aligning sequences according to the branching order in an initial phylogenetic tree. The results obtained using the package appear to be as sensitive as those from any other available method.
Desmond G. Higgins, P. M. Sharp
Comput. Appl. Biosci.1
1987 Interfacing similarity search software with the sequence retrieval system ACNUC
abstract
A method of interfacing sequence similarity search software with the fast sequence retrieval system ACNUC is described. The method is written in FORTRAN 77 and is straightforward to implement because no text-processing code is required--a minimum of 12 extra lines of FORTRAN provided the interface for most applications. The method is also efficient, since sequences are located by simple indexing techniques, with no linear searches of large database files necessary.
Desmond G. Higgins, Manolo Gouy
Comput. Appl. Biosci.1