VLDB 2026 Research / reviewers in the wild / expert
Ewan Birney
dblp:00/6856
· DBLP profile ↗
22ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-8314-8497ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Measurement and classification of bold-shy behaviours in medaka fishabstractMOTIVATION: Boldness-shyness is considered a fundamental axis of behavioural variation in humans and other species, with obvious adaptive causes and evolutionary implications. Besides an individual's own genetics, this phenotype is also affected by the genetic make-up of peers in the individual's social environment. To identify genetic determinants of variation along the bold-shy behavioural axis, a reliable experimental and analytical set-up able to highlight direct and indirect genetic effects is needed. RESULTS: We describe a custom assay designed to detect bold-shy behaviours in medaka fish, combining an open-field and novel-object component. We use this assay to explore direct and social genetic effects on the behaviours of 307 pairs of fish from five inbred medaka strains. Applying a hidden Markov model (HMM) to classify behavioural modes, we find that direct genetic effects influence the proportions of time the five strains spent in slow-moving states, explaining up to 29.7% of the variance in time spent in those states. We also found that an individual's behaviour is influenced by the genetics of its tank partner, explaining up to 8.64% of the variance in the time spent in slow-moving states. Our behavioural assay in combination with the HMM analysis is applicable to follow-up genetic linkage studies of genetic variants involved in direct behavioural effects and indirect social genetic effects. A suitable genetic resource for such studies, the Medaka Inbred Kiyosu-Karlsruhe (MIKK) panel has recently been established. AVAILABILITY AND IMPLEMENTATION: The code associated with this work is available on GitHub (https://github.com/birneylab/medaka_behaviour_pilot) and Software Heritage (swh: 1: dir: c9abec1c5d62d22e43c9e97d995c56261784d9ab). Experimental data have been uploaded to the EBI Bioimage Archive (https://doi.org/10.6019/S-BIAD1421). Saul Pierotti, Ian Brettell, Tomas W. Fitzgerald, Cathrin Herder, Narendar Aadepu, Christian Pylatiuk, Joachim Wittbrodt, Ewan Birney, Felix Loosli |
Bioinform. | 8 |
| 2025 | FlexLMM: a Nextflow linear mixed model framework for GWASabstractSUMMARY: Linear mixed models (LMMs) are a commonly used statistical approach in genome-wide association studies when population structure is present. However, naive permutations of the phenotype to empirically estimate the null distribution of a statistic of interest are not appropriate in the presence of population structure or covariates. This is because the samples are not exchangeable with each other under the null hypothesis, and because permuting the phenotypes breaks the relationship among those and eventual covariates. For this reason, we developed FlexLMM, a Nextflow pipeline that can perform appropriate permutations in LMMs while allowing for flexibility in the definition of the exact statistical model to be used. FlexLMM can set a significance threshold via permutations, thanks to a two-step process where the population structure is first regressed out, and only then are the permutations performed on the uncorrelated residuals. We envision this pipeline will be particularly useful for researchers working on multi-parental crosses among inbred lines of model organisms or farm animals and plants. AVAILABILITY AND IMPLEMENTATION: The source code and documentation for the FlexLMM is available at https://github.com/birneylab/flexlmm. Saul Pierotti, Tomas W. Fitzgerald, Ewan Birney |
Bioinform. | 3 |
| 2025 | Autoencoder-based phenotyping of ophthalmic images highlights genetic loci influencing retinal morphology and provides informative biomarkersabstractMOTIVATION: Genome-wide association studies (GWAS) have been remarkably successful in identifying associations between genetic variants and imaging-derived phenotypes. To date, the main focus of these analyses has been on established, clinically-used imaging features. We sought to investigate if deep learning approaches can detect more nuanced patterns of image variability. RESULTS: We used an autoencoder to represent retinal optical coherence tomography (OCT) images from 31 135 UK Biobank participants. For each subject, we obtained a 64-dimensional vector representing features of retinal structure. GWAS of these autoencoder-derived imaging parameters identified 118 statistically significant loci; 41 of these associations were also significant in a replication study. These loci encompassed variants previously linked with retinal thickness measurements, ophthalmic disorders, and/or neurodegenerative conditions. Notably, the generated retinal phenotypes were found to contribute to predictive models for glaucoma and cardiovascular disorders. Overall, we demonstrate that self-supervised phenotyping of OCT images enhances the discoverability of genetic factors influencing retinal morphology and provides epidemiologically informative biomarkers. AVAILABILITY AND IMPLEMENTATION: Code and data links available at https://github.com/tf2/autoencoder-oct. Panagiotis I. Sergouniotis, Adam Diakite, Ewan Birney, Tomas W. Fitzgerald |
Bioinform. | 5 |
| 2024 | Benchmarking of computational methods for m6A profiling with Nanopore direct RNA sequencingabstractN6-methyladenosine (m6A) is the most abundant internal eukaryotic mRNA modification, and is involved in the regulation of various biological processes. Direct Nanopore sequencing of native RNA (dRNA-seq) emerged as a leading approach for its identification. Several software were published for m6A detection and there is a strong need for independent studies benchmarking their performance on data from different species, and against various reference datasets. Moreover, a computational workflow is needed to streamline the execution of tools whose installation and execution remains complicated. We developed NanOlympicsMod, a Nextflow pipeline exploiting containerized technology for comparing 14 tools for m6A detection on dRNA-seq data. NanOlympicsMod was tested on dRNA-seq data generated from in vitro (un)modified synthetic oligos. The m6A hits returned by each tool were compared to the m6A position known by design of the oligos. In addition, NanOlympicsMod was used on dRNA-seq datasets from wild-type and m6A-depleted yeast, mouse and human, and each tool's hits were compared to reference m6A sets generated by leading orthogonal methods. The performance of the tools markedly differed across datasets, and methods adopting different approaches showed different preferences in terms of precision and recall. Changing the stringency cut-offs allowed for tuning the precision-recall trade-off towards user preferences. Finally, we determined that precision and recall of tools are markedly influenced by sequencing depth, and that additional sequencing would likely reveal additional m6A sites. Thanks to the possibility of including novel tools, NanOlympicsMod will streamline the benchmarking of m6A detection tools on dRNA-seq data, improving future RNA modification characterization. Simone Maestri, Mattia Furlan, Logan Mulroney, Lucia Coscujuela Tarrero, Camilla Ugolini, Fabio Dalla Pozza, Tommaso Leonardi, Ewan Birney, Francesco Nicassio, Mattia Pelizzola |
Briefings Bioinform. | 8 |
| 2024 | FEHAT: efficient, large scale and automated heartbeat detection in Medaka fish embryosabstractSUMMARY: High-resolution imaging of model organisms allows the quantification of important physiological measurements. In the case of fish with transparent embryos, these videos can visualize key physiological processes, such as heartbeat. High throughput systems can provide enough measurements for the robust investigation of developmental processes as well as the impact of system perturbations on physiological state. However, few analytical schemes have been designed to handle thousands of high-resolution videos without the need for some level of human intervention. We developed a software package, named FEHAT, to provide a fully automated solution for the analytics of large numbers of heart rate imaging datasets obtained from developing Medaka fish embryos in 96-well plate format imaged on an Acquifer machine. FEHAT uses image segmentation to define regions of the embryo showing changes in pixel intensity over time, followed by the classification of the most likely position of the heart and Fourier Transformations to estimate the heart rate. Here, we describe some important features of the FEHAT software, showcasing its performance across a large set of medaka fish embryos and compare its performance to established, less automated solutions. FEHAT provides reliable heart rate estimates across a range of temperature-based perturbations and can be applied to tens of thousands of embryos without the need for any human intervention. AVAILABILITY AND IMPLEMENTATION: Data used in this manuscript will be made available on request. Marcio Soares Ferreira, Sebastian Stricker, Tomas W. Fitzgerald, Jack Monahan, Fanny Defranoux, Philip Watson, Bettina Welz, Omar Hammouda, Joachim Wittbrodt, Ewan Birney |
Bioinform. | 10 |
| 2019 | htsget: a protocol for securely streaming genomic dataabstractSummary: Standardized interfaces for efficiently accessing high-throughput sequencing data are a fundamental requirement for large-scale genomic data sharing. We have developed htsget, a protocol for secure, efficient and reliable access to sequencing read and variation data. We demonstrate four independent client and server implementations, and the results of a comprehensive interoperability demonstration. Availability and implementation: http://samtools.github.io/hts-specs/htsget.html. Supplementary information: Supplementary data are available at Bioinformatics online. Jerome Kelleher, Mike Lin, Carl H. Albach, Ewan Birney, Robert Davies, Marina Gourtovaia, David Glazer, Cristina Y. González, David K. Jackson, Aaron Kemp, John Marshall, Andrew Nowak, Alexander Senf, Jaime M. Tovar-Corona, Alexander Vikhorev, Thomas M. Keane |
Bioinform. | 4 |
| 2018 | PhenotypeSimulator: A comprehensive framework for simulating multi-trait, multi-locus genotype to phenotype relationshipsabstractMotivation: Simulation is a critical part of method development and assessment. With the increasing sophistication of multi-trait and multi-locus genetic analysis techniques, it is important that the community has flexible simulation tools to challenge and explore the properties of these methods. Results: We have developed PhenotypeSimulator, a comprehensive phenotype simulation scheme that can model multiple traits with multiple underlying genetic loci as well as complex covariate and observational noise structure. This package has been designed to work with many common genetic tools both for input and output. We describe the underlying components of this simulation tool and illustrate its use on an example dataset. Availability and implementation: PhenotypeSimulator is available as a well documented R/CRAN package and the code is available on github: https://github.com/HannahVMeyer/PhenotypeSimulator. Supplementary information: Supplementary data are available at Bioinformatics online. Hannah Verena Meyer, Ewan Birney |
Bioinform. | 2 |
| 2018 | ChromoTrace: Computational reconstruction of 3D chromosome configurations for super-resolution microscopyabstractThe 3D structure of chromatin plays a key role in genome function, including gene expression, DNA replication, chromosome segregation, and DNA repair. Furthermore the location of genomic loci within the nucleus, especially relative to each other and nuclear structures such as the nuclear envelope and nuclear bodies strongly correlates with aspects of function such as gene expression. Therefore, determining the 3D position of the 6 billion DNA base pairs in each of the 23 chromosomes inside the nucleus of a human cell is a central challenge of biology. Recent advances of super-resolution microscopy in principle enable the mapping of specific molecular features with nanometer precision inside cells. Combined with highly specific, sensitive and multiplexed fluorescence labeling of DNA sequences this opens up the possibility of mapping the 3D path of the genome sequence in situ. Here we develop computational methodologies to reconstruct the sequence configuration of all human chromosomes in the nucleus from a super-resolution image of a set of fluorescent in situ probes hybridized to the genome in a cell. To test our approach, we develop a method for the simulation of DNA in an idealized human nucleus. Our reconstruction method, ChromoTrace, uses suffix trees to assign a known linear ordering of in situ probes on the genome to an unknown set of 3D in-situ probe positions in the nucleus from super-resolved images using the known genomic probe spacing as a set of physical distance constraints between probes. We find that ChromoTrace can assign the 3D positions of the majority of loci with high accuracy and reasonable sensitivity to specific genome sequences. By simulating appropriate spatial resolution, label multiplexing and noise scenarios we assess our algorithms performance. Our study shows that it is feasible to achieve genome-wide reconstruction of the 3D DNA path based on super-resolution microscopy images. Carl Barton, Sandro Morganella, Øyvind Ødegård-Fougner, Stephanie Alexander, Jonas Ries, Tomas W. Fitzgerald, Jan Ellenberg, Ewan Birney |
PLoS Comput. Biol. | 8 |
| 2014 | The EBI RDF platform: linked open data for the life sciencesabstractMOTIVATION: Resource description framework (RDF) is an emerging technology for describing, publishing and linking life science data. As a major provider of bioinformatics data and services, the European Bioinformatics Institute (EBI) is committed to making data readily accessible to the community in ways that meet existing demand. The EBI RDF platform has been developed to meet an increasing demand to coordinate RDF activities across the institute and provides a new entry point to querying and exploring integrated resources available at the EBI. Simon Jupp, James Malone, Jerven T. Bolleman, Marco Brandizi, Mark Davies, Leyla Jael Castro, Anna Gaulton, Sebastien Gehant, Camille Laibe, Nicole Redaschi, Sarala M. Wimalaratne, Maria Jesus Martin, Nicolas Le Novère, Helen E. Parkinson, Ewan Birney, Andrew M. Jenkinson |
Bioinform. | 15 |
| 2012 | Oases: robust de novo RNA-seq assembly across the dynamic range of expression levelsabstractMOTIVATION: High-throughput sequencing has made the analysis of new model organisms more affordable. Although assembling a new genome can still be costly and difficult, it is possible to use RNA-seq to sequence mRNA. In the absence of a known genome, it is necessary to assemble these sequences de novo, taking into account possible alternative isoforms and the dynamic range of expression values. RESULTS: We present a software package named Oases designed to heuristically assemble RNA-seq reads in the absence of a reference genome, across a broad spectrum of expression values and in presence of alternative isoforms. It achieves this by using an array of hash lengths, a dynamic filtering of noise, a robust resolution of alternative splicing events and the efficient merging of multiple assemblies. It was tested on human and mouse RNA-seq data and is shown to improve significantly on the transABySS and Trinity de novo transcriptome assemblers. AVAILABILITY AND IMPLEMENTATION: Oases is freely available under the GPL license at www.ebi.ac.uk/~zerbino/oases/. Marcel H. Schulz, Daniel R. Zerbino, Martin Vingron, Ewan Birney |
Bioinform. | 4 |
| 2010 | A database and API for variation, dense genotyping and resequencing dataabstractBACKGROUND: Advances in sequencing and genotyping technologies are leading to the widespread availability of multi-species variation data, dense genotype data and large-scale resequencing projects. The 1000 Genomes Project and similar efforts in other species are challenging the methods previously used for storage and manipulation of such data necessitating the redesign of existing genome-wide bioinformatics resources. RESULTS: Ensembl has created a database and software library to support data storage, analysis and access to the existing and emerging variation data from large mammalian and vertebrate genomes. These tools scale to thousands of individual genome sequences and are integrated into the Ensembl infrastructure for genome annotation and visualisation. The database and software system is easily expanded to integrate both public and non-public data sources in the context of an Ensembl software installation and is already being used outside of the Ensembl project in a number of database and application environments. CONCLUSIONS: Ensembl's powerful, flexible and open source infrastructure for the management of variation, genotyping and resequencing data is freely available at http://www.ensembl.org. Daniel Rios, William M. McLaren, Yuan Chen 0007, Ewan Birney, Arne Stabenau, Paul Flicek, Fiona Cunningham |
BMC Bioinform. | 4 |
| 2009 | Sequence progressive alignment, a framework for practical large-scale probabilistic consistency alignmentabstractMOTIVATION: Multiple sequence alignment is a cornerstone of comparative genomics. Much work has been done to improve methods for this task, particularly for the alignment of small sequences, and especially for amino acid sequences. However, less work has been done in making promising methods that work on the small-scale practically for the alignment of much larger genomic sequences. RESULTS: We take the method of probabilistic consistency alignment and make it practical for the alignment of large genomic sequences. In so doing we develop a set of new technical methods, combined in a framework we term 'sequence progressive alignment', because it allows us to iteratively compute an alignment by passing over the input sequences from left to right. The result is that we massively decrease the memory consumption of the program relative to a naive implementation. The general engineering of the challenges faced in scaling such a computationally intensive process offer valuable lessons for planning related large-scale sequence analysis algorithms. We also further show the strong performance of Pecan using an extended analysis of ancient repeat alignments. Pecan is now one of the default alignment programs that has and is being used by a number of whole-genome comparative genomic projects. AVAILABILITY: The Pecan program is freely available at http://www.ebi.ac.uk/ approximately bjp/pecan/ Pecan whole genome alignments can be found in the Ensembl genome browser. Benedict Paten, Javier Herrero, Kathryn Beal, Ewan Birney |
Bioinform. | 4 |
| 2008 | Integrating biological data - the Distributed Annotation SystemabstractBACKGROUND: The Distributed Annotation System (DAS) is a widely adopted protocol for dynamically integrating a wide range of biological data from geographically diverse sources. DAS continues to expand its applicability and evolve in response to new challenges facing integrative bioinformatics. RESULTS: Here we describe the various infrastructure components of DAS and present a new extended version of the DAS specification. Version 1.53E incorporates several recent developments, including its extension to serve new data types and an ontology for protein features. CONCLUSION: Our extensions to the DAS protocol have facilitated the integration of new data types, and our improvements to the existing DAS infrastructure have addressed recent challenges. The steadily increasing numbers of available data sources demonstrates further adoption of the DAS protocol. Andrew M. Jenkinson, Mario Albrecht, Ewan Birney, Hagen Blankenburg, Thomas A. Down, Robert D. Finn, Henning Hermjakob, Tim J. P. Hubbard, Rafael C. Jiménez, Philip Jones, Andreas Kähäri, Eugene Kulesha, José R. Macías, Gabrielle A. Reeves, Andreas Prlic |
BMC Bioinform. | 3 |
| 2008 | Advanced Genomic Data MiningabstractBioMart Web InterfaceFirst we will focus on BioMart's Web interface (http://www.biomart.org)to illustrate how to join two different datasets: Reactome [13], a database of metabolic pathways, and UniProt [14], a catalogue of protein information.In this example, we need to obtain a catalogue of enzymes involved in carbohydrate metabolism in humans, as we are interested in a congenic disorder in this pathway.To ask this question without an integrated data mining tool, one would have to start with Reactome to find enzymes involved in reaction pathways in human and then compare those enzymes to a list of entries in UniProt.However, BioMart allows us to join the two databases.We can start our query by clicking on 'MartView' from the Web interface at http://www.biomart.org,and selecting the Reactome database.Now, select the reaction dataset.Filters applied will be simply 'Limit to Species' Homo sapiens.Attributes can be selected as ''Reaction name'' and ''Gene ENSEMBL ID''.At this stage, 2,432 entries meet our criteria (i.e.we have asked for all human reaction pathways in the Reactome database).Click on the 'count' button at the top to obtain this number.Next, we can enrich our search for enzymes in the UniProt database.This will require the 'linked' or secondary dataset.Follow this description, or view the tutorials for use of the linked database at http://www.ensembl.org/common/Workshops_Online?id = 117.Click on the second 'Dataset' option at the left of the page.Select 'UniProt proteomes' as the database.In this instance, we will add as a filter the Gene Ontology (GO) [15] term 'GO:0005975' (associated with carbohydrate metabolic processes); this will be under 'EXTER-NAL IDENTIFIERS', 'Limit to pro-teins…GO ID(s)' in the secondary dataset.Also select, under 'External references': 'Entries with EC ID(s)', to limit our query to enzymes only, and 'eukaryota' along with 'Homo sapiens' under 'SPE-CIES' (Species and Proteome Name, respectively).This will give a count of 257 in the secondary dataset.The genome location can be displayed in the output by choosing the following Attributes: ''Genome component name'' for the chromosome, ''Start Position'' and ''End Position'' for the coordinates.Click 'Results' for the table in Figure 2. Now you have a list of enzymes in UniProt involved in carbohydrate metabolism in humans. BioConductorBioConductor is open source software for the analysis of genomic data.It is based Xosé M. Fernández, Ewan Birney |
PLoS Comput. Biol. | 2 |
| 2007 | Optimising oligonucleotide array design for ChIP-on-chip
Fiona G. G. Nielsen, Stefan Gräf, Stefan Kurtz, Sergei Denissov, Roland Green, Ewan Birney, Paul Flicek, Martijn A. Huynen, Henk Stunnenberg |
BMC Bioinform. | 7 |
| 2005 | Gene finding in the chicken genomeabstractBACKGROUND: Despite the continuous production of genome sequence for a number of organisms, reliable, comprehensive, and cost effective gene prediction remains problematic. This is particularly true for genomes for which there is not a large collection of known gene sequences, such as the recently published chicken genome. We used the chicken sequence to test comparative and homology-based gene-finding methods followed by experimental validation as an effective genome annotation method. RESULTS: We performed experimental evaluation by RT-PCR of three different computational gene finders, Ensembl, SGP2 and TWINSCAN, applied to the chicken genome. A Venn diagram was computed and each component of it was evaluated. The results showed that de novo comparative methods can identify up to about 700 chicken genes with no previous evidence of expression, and can correctly extend about 40% of homology-based predictions at the 5' end. CONCLUSIONS: De novo comparative gene prediction followed by experimental verification is effective at enhancing the annotation of the newly sequenced genomes provided by standard homology-based methods. Eduardo Eyras, Alexandre Reymond, Robert Castelo, Jacqueline M. Bye, Francisco Camara, Paul Flicek, Elizabeth J. Huckle, Genis Parra, David D. Shteynberg, Carine Wyss, Jane Rogers, Stylianos E. Antonarakis, Ewan Birney, Roderic Guigó, Michael R. Brent |
BMC Bioinform. | 13 |
| 2005 | Automated generation of heuristics for biological sequence comparisonabstractBACKGROUND: Exhaustive methods of sequence alignment are accurate but slow, whereas heuristic approaches run quickly, but their complexity makes them more difficult to implement. We introduce bounded sparse dynamic programming (BSDP) to allow rapid approximation to exhaustive alignment. This is used within a framework whereby the alignment algorithms are described in terms of their underlying model, to allow automated development of efficient heuristic implementations which may be applied to a general set of sequence comparison problems. RESULTS: The speed and accuracy of this approach compares favourably with existing methods. Examples of its use in the context of genome annotation are given. CONCLUSIONS: This system allows rapid implementation of heuristics approximating to many complex alignment models, and has been incorporated into the freely available sequence alignment program, exonerate. Guy St. C. Slater, Ewan Birney |
BMC Bioinform. | 2 |
| 2004 | Biological database design and implementationabstractWe present our experience of building biological databases. Such databases have most aspects in common with other complex databases in other fields. We do not believe that biological data are that different from complex data in other fields. Our experience has led us to emphasise simplicity and conservative technology choices when building these databases. This is a short paper of advice that we hope is useful to people designing their own biological database. Ewan Birney, Michele E. Clamp |
Briefings Bioinform. | 1 |
| 2000 | InterPro-an integrated documentation resource for protein families, domains and functional sitesabstractMOTIVATION: InterPro is a new integrated documentation resource for protein families, domains and functional sites, developed initially as a means of rationalising the complementary efforts of the PROSITE, PRINTS, Pfam and ProDom database projects. RESULTS: Merged annotations from PRINTS, PROSITE and Pfam form the InterPro core. Each combined InterPro entry includes functional descriptions and literature references, and links are made back to the relevant parent database(s), allowing users to see at a glance whether a particular family or domain has associated patterns, profiles, fingerprints, etc. Merged and individual entries (i.e. those that have no counterpart in the companion resources) are assigned unique accession numbers. Release 1.2 of InterPro (June 2000) contains over 3000 entries, representing families, domains, repeats and sites of post-translational modification (PTMs) encoded by 6581 different regular expressions, profiles, fingerprints and Hidden Markov Models (HMMs). Each InterPro entry lists all the matches against SWISS-PROT and TrEMBL (more than 1000000 hits from 264333 different proteins out of 384572 in SWISS-PROT and TrEMBL). Rolf Apweiler, Terri K. Attwood, Amos Bairoch, Alex Bateman, Ewan Birney, Margaret Biswas, Philipp Bucher, Lorenzo Cerutti, Florence Corpet, Michael D. R. Croning, Richard Durbin, Laurent Falquet, Wolfgang Fleischmann, Jérôme Gouzy, Henning Hermjakob, Nicolas Hulo, Inge Jonassen, Daniel Kahn, Alexander Kanapin, Youla Karavidopoulou, Rodrigo Lopez, Beate Marx, Nicola J. Mulder, Thomas M. Oinn, Marco Pagni, Florence Servant, Christian J. A. Sigrist, Evgeny M. Zdobnov |
Bioinform. | 5 |
| 2000 | ProtEST: protein multiple sequence alignments from expressed sequence tagsabstractMOTIVATION: An automatic sequence searching method (ProtEST) is described which constructs multiple protein sequence alignments from protein sequences and translated expressed sequence tags (ESTs). ProtEST is more effective than a simple TBLASTN search of the query against the EST database, as the sequences are automatically clustered, assembled, made non-redundant, checked for sequence errors, translated into protein and then aligned and displayed. RESULTS: A ProtEST search found a non-redundant, translated, error- and length-corrected EST sequence for > 58% of sequences when single sequences from 1407 Pfam-A seed alignments were used as the probe. The average family size of the resulting alignments of translated EST sequences contained > 10 sequences. In a cross-validated test of protein secondary structure prediction, alignments from the new procedure led to an improvement of 3.4% average Q3 prediction accuracy over single sequences. AVAILABILITY: The ProtEST method is available as an Internet World Wide Web service http://barton.ebi.ac.uk/servers/protest.html+ ++ The Wise2 package for protein and genomic comparisons and the ProtESTWise script can be found at http://www.sanger.ac.uk/Software/Wise2 CONTACT: [email protected] James A. Cuff, Ewan Birney, Michele E. Clamp, Geoffrey J. Barton |
Bioinform. | 2 |
| 1998 | SPEM: a parser for EMBL style flat file database entriesabstractSUMMARY: We present a set of Perl modules for the flexible and robust parsing and editing of EMBL/SWISS-PROT databases. AVAILABILITY: The Web page at http://www.sanger.ac. uk/Software/PerlModule/ provides information about downloading the SPEM and PrEMBL modules, and provides links to documentation and example code. Matthew R. Pocock, Tim J. P. Hubbard, Ewan Birney |
Bioinform. | 3 |
| 1997 | Dynamite: A Flexible Code Generating Language for Dynamic Programming Methods Used in Sequence Comparison
Ewan Birney, Richard Durbin |
ISMB | 1 |