EDBT 2026 Demo / reviewers in the wild / expert
Anton J. Enright
dblp:80/4508
· DBLP profile ↗
15ranked-venue papers
2as first author
2since 2021 · last 2026
0000-0002-6090-3100ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
11 papers |
Bioinformatics and computational biology · 100% |
Topics — the 21 heaviest of 22, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › transcriptomics › transcriptome sequencing
small RNA sequencing |
0.2 | 1 | 2015 | Chimira: analysis of small RNA sequencing data and microRNA modifications · Bioinform. 2015 |
Bioinformatics and computational biology › sequence analysis › motif discovery
regulatory motif discovery |
0.1 | 1 | 2010 | iMotifs: an integrated sequence motif visualization and analysis environment · Bioinform. 2010 |
Bioinformatics and computational biology › sequence analysis
sequence motif analysis |
0.1 | 1 | 2010 | iMotifs: an integrated sequence motif visualization and analysis environment · Bioinform. 2010 |
Bioinformatics and computational biology › gene regulation › transcription factor binding
transcription factor binding site |
0.1 | 1 | 2010 | iMotifs: an integrated sequence motif visualization and analysis environment · Bioinform. 2010 |
Bioinformatics and computational biology
comparative genomics |
0.1 | 2 | 2005 | CoGenT++: an extensive and extensible data environment for computational genomics · Bioinform. 2005 Transcription-associated protein families are primarily taxon-specific · Bioinform. 2001 |
Bioinformatics and computational biology
functional genomics |
0.1 | 1 | 2005 | CoGenT++: an extensive and extensible data environment for computational genomics · Bioinform. 2005 |
Bioinformatics and computational biology › biological database
sequence database |
0.1 | 1 | 2005 | MagicMatch - cross-referencing sequence identifiers across databases · Bioinform. 2005 |
Bioinformatics and computational biology › genome annotation
annotation quality assessment |
0.0 | 1 | 2003 | Evaluation of annotation strategies using an entire genome sequence · Bioinform. 2003 |
Bioinformatics and computational biology › genome annotation
functional annotation |
0.0 | 1 | 2003 | Evaluation of annotation strategies using an entire genome sequence · Bioinform. 2003 |
Bioinformatics and computational biology
genome annotation |
0.0 | 1 | 2003 | Evaluation of annotation strategies using an entire genome sequence · Bioinform. 2003 |
Bioinformatics and computational biology › genomics › genomic data management
genome database |
0.0 | 1 | 2003 | COmplete GENome Tracking (COGENT): A Flexible Data Environment for Computational Genomics · Bioinform. 2003 |
Bioinformatics and computational biology
genomics |
0.0 | 1 | 2003 | COmplete GENome Tracking (COGENT): A Flexible Data Environment for Computational Genomics · Bioinform. 2003 |
Bioinformatics and computational biology
gene expression analysis |
0.0 | 1 | 2010 | SylArray: a web server for automated detection of miRNA effects from expression data · Bioinform. 2010 |
Bioinformatics and computational biology
biological data visualization |
0.0 | 1 | 2001 | BioLayout-an automatic graph layout algorithm for similarity visualization · Bioinform. 2001 |
Bioinformatics and computational biology › protein sequence analysis › protein sequence annotation
low complexity region analysis |
0.0 | 1 | 2000 | CAST: an iterative algorithm for the complexity analysis of sequence tracts · Bioinform. 2000 |
Bioinformatics and computational biology › protein structure analysis
protein domain identification |
0.0 | 1 | 2000 | GeneRAGE: a robust algorithm for sequence clustering and domain detection · Bioinform. 2000 |
Bioinformatics and computational biology › protein function prediction › protein classification
protein family classification |
0.0 | 1 | 2000 | GeneRAGE: a robust algorithm for sequence clustering and domain detection · Bioinform. 2000 |
Bioinformatics and computational biology › protein sequence analysis
protein sequence clustering |
0.0 | 1 | 2000 | GeneRAGE: a robust algorithm for sequence clustering and domain detection · Bioinform. 2000 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 2000 | CAST: an iterative algorithm for the complexity analysis of sequence tracts · Bioinform. 2000 |
Bioinformatics and computational biology › genome annotation
automatic annotation |
0.0 | 1 | 2003 | Evaluation of annotation strategies using an entire genome sequence · Bioinform. 2003 |
Bioinformatics and computational biology
phylogenetics |
0.0 | 1 | 2001 | Transcription-associated protein families are primarily taxon-specific · Bioinform. 2001 |
Methods — techniques the papers use, named apart from their topics
trimming · 0.2statistical analysis · 0.2sequence mapping · 0.2word enrichment · 0.1sylamer algorithm · 0.1NestedMICA · 0.1smith-waterman alignment · 0.1phylogenetic profiling · 0.1hashing · 0.1all-against-all similarity · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PymiRa: A rapid and accurate classification tool for small non-coding RNAs, including microRNAsabstractSmall non-coding RNAs (sncRNA; < 200 nucleotide length) are of increasing research interest due to their key regulatory roles in a host of fundamental biological processes. For example, microRNAs (miRNAs), a specific class of sncRNAs, regulate gene expression through messenger RNA (mRNA) interactions, and their dysregulation is associated with disease. Classifying sncRNAs is an important bioinformatic task in small RNA-sequencing pipelines. Here we have developed an aligner called PymiRa, written in Python, to identify and quantify miRNAs from FASTA/FASTQ sequencing files. Unlike other approaches, PymiRa utilises a Burrows-Wheeler algorithm to align an input file against a reference hairpin precursor FASTA file derived from miRBase, the online miRNA registry, permitting up to two mismatches at the 3' end of a read. Previous tools used either a Burrows-Wheeler genome alignment or dynamic programming alignment to precursors; we demonstrate that combining both approaches yields improved results and efficiency. Importantly, the PymiRa aligner accounts for 3' post-transcriptional modifications to miRNAs that typically occur. PymiRa is a fast, accurate, and publicly accessible aligner available via GitHub and/or a webserver for sncRNA identification, including miRNAs, enabling accurate counts to be produced as part of a small RNA-sequencing pipeline. PymiRa will undergo relevant revisions over time e.g., with miRBase version updates. The PymiRa aligner will facilitate a deeper biological understanding of the landscape of sncRNA expression in normal physiological conditions and their dysregulation in disease states, including cancer. Zachary G. L. Scurlock, Cinzia G. Scarpini, Nicholas Coleman, Matthew J. Murray, Anton J. Enright |
PLoS Comput. Biol. | 5 |
| 2023 | CGG toolkit: Software components for computational genomicsabstractPublic-domain availability for bioinformatics software resources is a key requirement that ensures long-term permanence and methodological reproducibility for research and development across the life sciences. These issues are particularly critical for widely used, efficient, and well-proven methods, especially those developed in research settings that often face funding discontinuities. We re-launch a range of established software components for computational genomics, as legacy version 1.0.1, suitable for sequence matching, masking, searching, clustering and visualization for protein family discovery, annotation and functional characterization on a genome scale. These applications are made available online as open source and include MagicMatch, GeneCAST, support scripts for CoGenT-like sequence collections, GeneRAGE and DifFuse, supported by centrally administered bioinformatics infrastructure funding. The toolkit may also be conceived as a flexible genome comparison software pipeline that supports research in this domain. We illustrate basic use by examples and pictorial representations of the registered tools, which are further described with appropriate documentation files in the corresponding GitHub release. Dimitrios Vasileiou, Christos Karapiperis, Ismini Baltsavia, Anastasia Chasapi, Dag G. Ahrén, Paul J. Janssen, Vasilis J. Promponas, Anton J. Enright, Christos A. Ouzounis |
PLoS Comput. Biol. | 9 |
| 2015 | Chimira: analysis of small RNA sequencing data and microRNA modificationsabstractUNLABELLED: Chimira is a web-based system for microRNA (miRNA) analysis from small RNA-Seq data. Sequences are automatically cleaned, trimmed, size selected and mapped directly to miRNA hairpin sequences. This generates count-based miRNA expression data for subsequent statistical analysis. Moreover, it is capable of identifying epi-transcriptomic modifications in the input sequences. Supported modification types include multiple types of 3'-modifications (e.g. uridylation, adenylation), 5'-modifications and also internal modifications or variation (ADAR editing or single nucleotide polymorphisms). Besides cleaning and mapping of input sequences to miRNAs, Chimira provides a simple and intuitive set of tools for the analysis and interpretation of the results (see also Supplementary Material). These allow the visual study of the differential expression between two specific samples or sets of samples, the identification of the most highly expressed miRNAs within sample pairs (or sets of samples) and also the projection of the modification profile for specific miRNAs across all samples. Other tools have already been published in the past for various types of small RNA-Seq analysis, such as UEA workbench, seqBuster, MAGI, OASIS and CAP-miRSeq, CPSS for modifications identification. A comprehensive comparison of Chimira with each of these tools is provided in the Supplementary Material. Chimira outperforms all of these tools in total execution speed and aims to facilitate simple, fast and reliable analysis of small RNA-Seq data allowing also, for the first time, identification of global microRNA modification profiles in a simple intuitive interface. AVAILABILITY AND IMPLEMENTATION: Chimira has been developed as a web application and it is accessible here: http://www.ebi.ac.uk/research/enright/software/chimira. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Dimitrios M. Vitsios, Anton J. Enright |
Bioinform. | 2 |
| 2010 | SylArray: a web server for automated detection of miRNA effects from expression dataabstractUNLABELLED: A useful step for understanding the function of microRNAs (miRNA) or siRNAs is the detection of their effects on genome-wide expression profiles. Typically, approaches look for enrichment of words in the 3(')UTR sequences of the most deregulated genes. A number of tools are available for this purpose, but they require either in-depth computational knowledge, filtered 3(')UTR sequences for the genome of interest, or a set of genes acquired through an arbitrary expression cutoff. To this end, we have developed SylArray; a web-based resource designed for the analysis of large-scale expression datasets. It simply requires the user to submit a sorted list of genes from an expression experiment. SylArray utilizes curated sets of 3(')UTRs to attach sequences to these genes and then applies the Sylamer algorithm for detection of miRNA or siRNA signatures in those sequences. An intuitive system for visualization and interpretation of the small RNA signatures is included. AVAILABILITY: SylArray is written in Perl-CGI, Perl and Java and also uses the R statistical package. The source-code, database and web resource are freely available under GNU Public License (GPL). The web server is freely accessible at http://www.ebi.ac.uk/enright/sylarray. Nenad Bartonicek, Anton J. Enright |
Bioinform. | 2 |
| 2010 | iMotifs: an integrated sequence motif visualization and analysis environmentabstractMOTIVATION: Short sequence motifs are an important class of models in molecular biology, used most commonly for describing transcription factor binding site specificity patterns. High-throughput methods have been recently developed for detecting regulatory factor binding sites in vivo and in vitro and consequently high-quality binding site motif data are becoming available for increasing number of organisms and regulatory factors. Development of intuitive tools for the study of sequence motifs is therefore important. iMotifs is a graphical motif analysis environment that allows visualization of annotated sequence motifs and scored motif hits in sequences. It also offers motif inference with the sensitive NestedMICA algorithm, as well as overrepresentation and pairwise motif matching capabilities. All of the analysis functionality is provided without the need to convert between file formats or learn different command line interfaces. The application includes a bundled and graphically integrated version of the NestedMICA motif inference suite that has no outside dependencies. Problems associated with local deployment of software are therefore avoided. AVAILABILITY: iMotifs is licensed with the GNU Lesser General Public License v2.0 (LGPL 2.0). The software and its source is available at http://wiki.github.com/mz2/imotifs and can be run on Mac OS X Leopard (Intel/PowerPC). We also provide a cross-platform (Linux, OS X, Windows) LGPL 2.0 licensed library libxms for the Perl, Ruby, R and Objective-C programming languages for input and output of XMS formatted annotated sequence motif set files. CONTACT: [email protected]; [email protected]. Matias Piipari, Thomas A. Down, Harpreet Kaur Saini, Anton J. Enright, Tim J. P. Hubbard |
Bioinform. | 4 |
| 2010 | MapMi: automated mapping of microRNA lociabstractBACKGROUND: A large effort to discover microRNAs (miRNAs) has been under way. Currently miRBase is their primary repository, providing annotations of primary sequences, precursors and probable genomic loci. In many cases miRNAs are identical or very similar between related (or in some cases more distant) species. However, miRBase focuses on those species for which miRNAs have been directly confirmed. Secondly, specific miRNAs or their loci are sometimes not annotated even in well-covered species. We sought to address this problem by developing a computational system for automated mapping of miRNAs within and across species. Given the sequence of a known miRNA in one species it is relatively straightforward to determine likely loci of that miRNA in other species. Our primary goal is not the discovery of novel miRNAs but the mapping of validated miRNAs in one species to their most likely orthologues in other species. RESULTS: We present MapMi, a computational system for automated miRNA mapping across and within species. This method has a sensitivity of 92.20% and a specificity of 97.73%. Using the latest release (v14) of miRBase, we obtained 10,944 unannotated potential miRNAs when MapMi was applied to all 21 species in Ensembl Metazoa release 2 and 46 species from Ensembl release 55. CONCLUSIONS: The pipeline and an associated web-server for mapping miRNAs are freely available on http://www.ebi.ac.uk/enright-srv/MapMi/. In addition precomputed miRNA mappings of miRBase miRNAs across a large number of species are provided. José Afonso Guerra-Assunção, Anton J. Enright |
BMC Bioinform. | 2 |
| 2007 | Construction, Visualisation, and Clustering of Transcription Networks from Microarray Expression DataabstractNetwork analysis transcends conventional pairwise approaches to data analysis as the context of components in a network graph can be taken into account. Such approaches are increasingly being applied to genomics data, where functional linkages are used to connect genes or proteins. However, while microarray gene expression datasets are now abundant and of high quality, few approaches have been developed for analysis of such data in a network context. We present a novel approach for 3-D visualisation and analysis of transcriptional networks generated from microarray data. These networks consist of nodes representing transcripts connected by virtue of their expression profile similarity across multiple conditions. Analysing genome-wide gene transcription across 61 mouse tissues, we describe the unusual topography of the large and highly structured networks produced, and demonstrate how they can be used to visualise, cluster, and mine large datasets. This approach is fast, intuitive, and versatile, and allows the identification of biological relationships that may be missed by conventional analysis techniques. This work has been implemented in a freely available open-source application named BioLayout Express(3D). Tom C. Freeman, Leon Goldovsky, Markus Brosch, Stijn van Dongen, Pierre Mazière, Russell J. Grocock, Shiri Freilich, Janet M. Thornton, Anton J. Enright |
PLoS Comput. Biol. | 9 |
| 2005 | CoGenT++: an extensive and extensible data environment for computational genomicsabstractMOTIVATION: CoGenT++ is a data environment for computational research in comparative and functional genomics, designed to address issues of consistency, reproducibility, scalability and accessibility. DESCRIPTION: CoGenT++ facilitates the re-distribution of all fully sequenced and published genomes, storing information about species, gene names and protein sequences. We describe our scalable implementation of ProXSim, a continually updated all-against-all similarity database, which stores pairwise relationships between all genome sequences. Based on these similarities, derived databases are generated for gene fusions--AllFuse, putative orthologs--OFAM, protein families--TRIBES, phylogenetic profiles--ProfUse and phylogenetic trees. Extensions based on the CoGenT++ environment include disease gene prediction, pattern discovery, automated domain detection, genome annotation and ancestral reconstruction. CONCLUSION: CoGenT++ provides a comprehensive environment for computational genomics, accessible primarily for large-scale analyses as well as manual browsing. Leon Goldovsky, Paul J. Janssen, Dag G. Ahrén, Benjamin Audit, Ildefonso Cases, Nikos Darzentas, Anton J. Enright, Núria López-Bigas, José M. Peregrín-Alvarez, Mike Smith 0001, Sophia Tsoka, Victor Kunin, Christos A. Ouzounis |
Bioinform. | 7 |
| 2005 | MagicMatch - cross-referencing sequence identifiers across databasesabstractMotivation: At present, mapping of sequence identifiers across databases is a daunting, time-consuming and computationally expensive process, usually achieved by sequence similarity searches with strict threshold values. Summary: We present a rapid and efficient method to map sequence identifiers across databases. The method uses the MD5 checksum algorithm for message integrity to generate sequence fingerprints and uses these fingerprints as hash strings to map sequences across databases. The program, called MagicMatch, is able to cross-link any of the major sequence databases within a few seconds on a modest desktop computer. Availability: MagicMatch is available at the following URL (http://cgg.ebi.ac.uk/services/magicmatch/), including an interactive service for major databases and binary downloads for widely used platforms. Contact: [email protected] Mike Smith 0001, Victor Kunin, Leon Goldovsky, Anton J. Enright, Christos A. Ouzounis |
Bioinform. | 4 |
| 2003 | Evaluation of annotation strategies using an entire genome sequenceabstractAbstract Motivation: Genome-wide functional annotation either by manual or automatic means has raised considerable concerns regarding the accuracy of assignments and the reproducibility of methodologies. In addition, a performance evaluation of automated systems that attempt to tackle sequence analyses rapidly and reproducibly is generally missing. In order to quantify the accuracy and reproducibility of function assignments on a genome-wide scale, we have re-annotated the entire genome sequence of Chlamydia trachomatis (serovar D), in a collaborative manner. Results: We have encoded all annotations in a structured format to allow further comparison and data exchange and have used a scale that records the different levels of potential annotation errors according to their propensity to propagate in the database due to transitive function assignments. We conclude that genome annotation may entail a considerable amount of errors, ranging from simple typographical errors to complex sequence analysis problems. The most surprising result of this comparative study is that automatic systems might perform as well as the teams of experts annotating genome sequences. Availability and supplementary information: http://www.ebi.ac.uk/research/cgg/annotation/cteval/ Contact: [email protected] * To whom correspondence should be addressed. † INA-EKETA, GR-57001 Thessaloniki, Greece ‡ Computational Biology Center, Memorial Sloan-Kettering Cancer Center, New York, NY 10021, USA § Aetion Technologies LLC, Worthington, OH 43085, USA ¶ Institut Curie, F-75248 Paris, France ∥ CNRS, UMR6543, F-06108 Nice, France ** Alma Bioinformatics, E-28760 Madrid, Spain †† Cap Gemini Ernst & Young, London SW1X 7LX, UK ‡‡ Univ. of Rome ‘La Sapienza’, I-00185 Rome, Italy §§ MWG-Biotech AG, Ebersberg, D-85560 Berlin, Germany ¶¶ Wellcome Trust Biocentre, Univ. of Dundee, Dundee DD1 5HN, UK Sophia Tsoka, Miguel A. Andrade-Navarro, Anton J. Enright, Mark Carroll, Patrick Poullet, Vasilis J. Promponas, Theodore Liakopoulos, Giorgos Palaios, Claude Pasquier, Stavros J. Hamodrakas, Javier Tamames, Asutosh T. Yagnik, Anna Tramontano, Damien Devos, Christian Blaschke, Alfonso Valencia, David Brett, David M. A. Martin, Christophe Leroy, Isidore Rigoutsos, Chris Sander, Christos A. Ouzounis |
Bioinform. | 4 |
| 2003 | COmplete GENome Tracking (COGENT): A Flexible Data Environment for Computational GenomicsabstractAbstract Summary: We present a database of fully sequenced and published genomes to facilitate the re-distribution of data and ensure reproducibility of results in the field of computational genomics. For its design we have implemented an extremely simple yet powerful schema to allow linking of genome sequence data to other resources. Availability: http://maine.ebi.ac.uk:8000/services/cogent/ Contact: [email protected] * To whom correspondence should be addressed. † The authors wish it to be known that, in their opinion, the first two authors should be regarded as joint First Authors. ‡ Present Address: External Services Group, EMBL-EBI. Paul J. Janssen, Anton J. Enright, Benjamin Audit, Ildefonso Cases, Leon Goldovsky, Nicola Harte, Victor Kunin, Christos A. Ouzounis |
Bioinform. | 2 |
| 2001 | Transcription-associated protein families are primarily taxon-specificabstractThe mechanisms controlling gene regulation appear to be fundamentally different in eukaryotes and prokaryotes (Struhl (1999) CELL, 98, 1-4). To investigate this diversity further, we have analysed the distribution of all known transcription-associated proteins (TAPs), as reflected by sequence database annotations. Our results for the primary phylogenetic domains (Archaea, Bacteria and Eukaryota) show that TAP families are mostly taxon-specific and very few transcriptional regulators are common across these domains. Richard M. R. Coulson, Anton J. Enright, Christos A. Ouzounis |
Bioinform. | 2 |
| 2001 | BioLayout-an automatic graph layout algorithm for similarity visualizationabstractUNLABELLED: Graph layout is extensively used in the field of mathematics and computer science, however these ideas and methods have not been extended in a general fashion to the construction of graphs for biological data. To this end, we have implemented a version of the Fruchterman Rheingold graph layout algorithm, extensively modified for the purpose of similarity analysis in biology. This algorithm rapidly and effectively generates clear two (2D) or three-dimensional (3D) graphs representing similarity relationships such as protein sequence similarity. The implementation of the algorithm is general and applicable to most types of similarity information for biological data. AVAILABILITY: BioLayout is available for most UNIX platforms at the following web-site: http://www.ebi.ac.uk/research/cgg/services/layout. Anton J. Enright, Christos A. Ouzounis |
Bioinform. | 1 |
| 2000 | GeneRAGE: a robust algorithm for sequence clustering and domain detectionabstractMOTIVATION: Efficient, accurate and automatic clustering of large protein sequence datasets, such as complete proteomes, into families, according to sequence similarity. Detection and correction of false positive and negative relationships with subsequent detection and resolution of multi-domain proteins. RESULTS: A new algorithm for the automatic clustering of protein sequence datasets has been developed. This algorithm represents all similarity relationships within the dataset in a binary matrix. Removal of false positives is achieved through subsequent symmetrification of the matrix using a Smith-Waterman dynamic programming alignment algorithm. Detection of multi-domain protein families and further false positive relationships within the symmetrical matrix is achieved through iterative processing of matrix elements with successive rounds of Smith-Waterman dynamic programming alignments. Recursive single-linkage clustering of the corrected matrix allows efficient and accurate family representation for each protein in the dataset. Initial clusters containing multi-domain families, are split into their constituent clusters using the information obtained by the multi-domain detection step. This algorithm can hence quickly and accurately cluster large protein datasets into families. Problems due to the presence of multi-domain proteins are minimized, allowing more precise clustering information to be obtained automatically. AVAILABILITY: GeneRAGE (version 1.0) executable binaries for most platforms may be obtained from the authors on request. The system is available to academic users free of charge under license. Anton J. Enright, Christos A. Ouzounis |
Bioinform. | 1 |
| 2000 | CAST: an iterative algorithm for the complexity analysis of sequence tractsabstractMOTIVATION: Sensitive detection and masking of low-complexity regions in protein sequences. Filtered sequences can be used in sequence comparison without the risk of matching compositionally biased regions. The main advantage of the method over similar approaches is the selective masking of single residue types without affecting other, possibly important, regions. RESULTS: A novel algorithm for low-complexity region detection and selective masking. The algorithm is based on multiple-pass Smith-Waterman comparison of the query sequence against twenty homopolymers with infinite gap penalties. The output of the algorithm is both the masked query sequence for further analysis, e.g. database searches, as well as the regions of low complexity. The detection of low-complexity regions is highly specific for single residue types. It is shown that this approach is sufficient for masking database query sequences without generating false positives. The algorithm is benchmarked against widely available algorithms using the 210 genes of Plasmodium falciparum chromosome 2, a dataset known to contain a large number of low-complexity regions. AVAILABILITY: CAST (version 1.0) executable binaries are available to academic users free of charge under license. Web site entry point, server and additional material: http://www.ebi.ac.uk/research/cgg/services/cast/ Vasilis J. Promponas, Anton J. Enright, Sophia Tsoka, David P. Kreil, Christophe Leroy, Stavros J. Hamodrakas, Chris Sander, Christos A. Ouzounis |
Bioinform. | 2 |