EDBT 2026 Demo / reviewers in the wild / expert
Thomas D. Otto
dblp:24/1282
· DBLP profile ↗
15ranked-venue papers
4as first author
6since 2021 · last 2025
0000-0002-1246-7404ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TrAGEDy - trajectory alignment of gene expression dynamicsabstractMOTIVATION: Single-cell transcriptomics sequencing is used to compare different biological processes. However, often, those processes are asymmetric which are difficult to integrate. Current approaches often rely on integrating samples from each condition before either cluster-based comparisons or analysis of an inferred shared trajectory. RESULTS: We present Trajectory Alignment of Gene Expression Dynamics (TrAGEDy), which allows the alignment of independent trajectories to avoid the need for error-prone integration steps. Across simulated datasets, TrAGEDy returns the correct underlying alignment of the datasets, outperforming current tools which fail to capture the complexity of asymmetric alignments. When applied to real datasets, TrAGEDy captures more biologically relevant genes and processes, which other differential expression methods fail to detect when looking at the developments of T cells and the bloodstream forms of Trypanosoma brucei when affected by genetic knockouts. AVAILABILITY AND IMPLEMENTATION: TrAGEDy is freely available at https://github.com/No2Ross/TrAGEDy, and implemented in R. Ross F. Laidlaw, Emma M. Briggs, Keith R. Matthews, Amir Madany Mamlouk, Richard McCulloch, Thomas D. Otto |
Bioinform. | 6 |
| 2025 | Exploring homology detection via k-means clustering of proteins embedded with a large language modelabstractMOTIVATION: Inferring protein homology from sequence information is essential for understanding species evolution and enabling functional annotation transfer. Besides similarity-based methods, several machine learning approaches have been developed using various ways of representing protein data. RESULTS: Here, we represent proteins with a biologically oriented large language model and apply k-means clustering to the embedded data to extract homology relationships. Although our approach lacks the sensitivity of other tools, we obtain better precision for the detection of n:m orthologs. Furthermore, we successfully reconstruct full orthologous groups from scratch, highlighting the growing potential of using large language models in combination with clustering algorithms for the analysis of protein data. AVAILABILITY AND IMPLEMENTATION: Datasets are available on OrthoMCL-DB as indicated in the Methods. Source code is available on GitHub at https://github.com/ThomasGTHB/OrthoLM and Zenodo at https://doi.org/10.5281/zenodo.16640170. Thomas Minotto, Antoine Claessens, Thomas D. Otto |
Bioinform. | 3 |
| 2023 | From contigs towards chromosomes: automatic improvement of long read assemblies (ILRA)abstractRecent advances in long read technologies not only enable large consortia to aim to sequence all eukaryotes on Earth, but they also allow individual laboratories to sequence their species of interest with relatively low investment. Long read technologies embody the promise of overcoming scaffolding problems associated with repeats and low complexity sequences, but the number of contigs often far exceeds the number of chromosomes and they may contain many insertion and deletion errors around homopolymer tracts. To overcome these issues, we have implemented the ILRA pipeline to correct long read-based assemblies. Contigs are first reordered, renamed, merged, circularized, or filtered if erroneous or contaminated. Illumina short reads are used subsequently to correct homopolymer errors. We successfully tested our approach by improving the genome sequences of Homo sapiens, Trypanosoma brucei, and Leptosphaeria spp., and by generating four novel Plasmodium falciparum assemblies from field samples. We found that correcting homopolymer tracts reduced the number of genes incorrectly annotated as pseudogenes, but an iterative approach seems to be required to correct more sequencing errors. In summary, we describe and benchmark the performance of our new tool, which improved the quality of novel long read assemblies up to 1 Gbp. The pipeline is available at GitHub: https://github.com/ThomasDOtto/ILRA. Susanne Reimering, Juan David Escobar-Prieto, Nicolas M. B. Brancucci, Diego F. Echeverry, Abdirahman I Abdi, Matthias Marti, Elena Gómez-Díaz, Thomas D. Otto |
Briefings Bioinform. | 9 |
| 2023 | peaks2utr: a robust Python tool for the annotation of 3′ UTRsabstractSUMMARY: Annotation of nonmodel organisms is an open problem, especially the detection of untranslated regions (UTRs). Correct annotation of UTRs is crucial in transcriptomic analysis to accurately capture the expression of each gene yet is mostly overlooked in annotation pipelines. Here we present peaks2utr, an easy-to-use Python command line tool that uses the UTR enrichment of single-cell technologies, such as 10× Chromium, to accurately annotate 3' UTRs for a given canonical annotation. AVAILABILITY AND IMPLEMENTATION: peaks2utr is implemented in Python 3 (≥3.8). It is available via PyPI at https://pypi.org/project/peaks2utr and GitHub at https://github.com/haessar/peaks2utr. It is licensed under GNU GPLv3. William Haese-Hill, Kathryn Crouch, Thomas D. Otto |
Bioinform. | 3 |
| 2022 | Varia: a tool for prediction, analysis and visualisation of variable genesabstractBACKGROUND: Parasites use polymorphic gene families to evade the immune system or interact with the host. Assessing the diversity and expression of such gene families in pathogens can inform on the repertoire or host interaction phenotypes of clinical relevance. However, obtaining the sequences and quantifying their expression is a challenge. In Plasmodium falciparum, the highly polymorphic var genes encode the major virulence protein, PfEMP1, which bind a range of human receptors through varying combinations of DBL and CIDR domains. Here we present a tool, Varia, to predict near full-length gene sequences and domain compositions of query genes from database genes sharing short sequence tags. Varia generates output through two complementary pipelines. Varia_VIP returns all putative gene sequences and domain compositions of the query gene from any partial sequence provided, thereby enabling experimental validation of specific genes of interest and detailed assessment of their putative domain structure. Varia_GEM accommodates rapid profiling of var gene expression in complex patient samples from DBLα expression sequence tags (EST), by computing a sample overall transcript profile stratified by PfEMP1 domain types. RESULTS: Varia_VIP was tested querying sequence tags from all DBL domain types using different search criteria. On average 92% of query tags had one or more 99% identical database hits, resulting in the full-length query gene sequence being identified (> 99% identical DNA > 80% of query gene) among the five most prominent database hits, for ~ 33% of the query genes. Optimized Varia_GEM settings allowed correct prediction of > 90% of domains placed among the four most N-terminal domains, including the DBLα domain, and > 70% of C-terminal domains. With this accuracy, N-terminal domains could be predicted for > 80% of queries, whereas prediction rates of C-terminal domains dropped with the distance from the DBLα from 70 to 40%. CONCLUSION: Prediction of var sequence and domain composition is possible from short sequence tags. Varia can be used to guide experimental validation of PfEMP1 sequences of interest and conduct high-throughput analysis of var type expression in patient samples. Gavin Mackenzie, Rasmus W. Jensen, Thomas Lavstsen, Thomas D. Otto |
BMC Bioinform. | 4 |
| 2022 | Using topic modeling to detect cellular crosstalk in scRNA-seqabstractCell-cell interactions are vital for numerous biological processes including development, differentiation, and response to inflammation. Currently, most methods for studying interactions on scRNA-seq level are based on curated databases of ligands and receptors. While those methods are useful, they are limited to our current biological knowledge. Recent advances in single cell protocols have allowed for physically interacting cells to be captured, and as such we have the potential to study interactions in a complemantary way without relying on prior knowledge. We introduce a new method based on Latent Dirichlet Allocation (LDA) for detecting genes that change as a result of interaction. We apply our method to synthetic datasets to demonstrate its ability to detect genes that change in an interacting population compared to a reference population. Next, we apply our approach to two datasets of physically interacting cells to identify the genes that change as a result of interaction, examples include adhesion and co-stimulatory molecules which confirm physical interaction between cells. For each dataset we produce a ranking of genes that are changing in subpopulations of the interacting cells. In addition to the genes discussed in the original publications, we highlight further candidates for interaction in the top 100 and 300 ranked genes. Lastly, we apply our method to a dataset generated by a standard droplet-based protocol not designed to capture interacting cells, and discuss its suitability for analysing interactions. We present a method that streamlines detection of interactions and does not require prior clustering and generation of synthetic reference profiles to detect changes in expression. Alexandrina Pancheva, Helen Wheadon, Simon Rogers, Thomas D. Otto |
PLoS Comput. Biol. | 4 |
| 2015 | IVA: accurate de novo assembly of RNA virus genomesabstractMOTIVATION: An accurate genome assembly from short read sequencing data is critical for downstream analysis, for example allowing investigation of variants within a sequenced population. However, assembling sequencing data from virus samples, especially RNA viruses, into a genome sequence is challenging due to the combination of viral population diversity and extremely uneven read depth caused by amplification bias in the inevitable reverse transcription and polymerase chain reaction amplification process of current methods. RESULTS: We developed a new de novo assembler called IVA (Iterative Virus Assembler) designed specifically for read pairs sequenced at highly variable depth from RNA virus samples. We tested IVA on datasets from 140 sequenced samples from human immunodeficiency virus-1 or influenza-virus-infected people and demonstrated that IVA outperforms all other virus de novo assemblers. AVAILABILITY AND IMPLEMENTATION: The software runs under Linux, has the GPLv3 licence and is freely available from http://sanger-pathogens.github.io/iva Martin Hunt, Astrid Gall, Swee Hoe Ong, Jacqui Brener, Bridget Ferns, Philip J. R. Goulder, Eleni Nastouli, Jacqueline A. Keane, Paul Kellam, Thomas D. Otto |
Bioinform. | 10 |
| 2013 | BamView: visualizing and interpretation of next-generation sequencing read alignmentsabstractUNLABELLED: So-called next-generation sequencing (NGS) has provided the ability to sequence on a massive scale at low cost, enabling biologists to perform powerful experiments and gain insight into biological processes. BamView has been developed to visualize and analyse sequence reads from NGS platforms, which have been aligned to a reference sequence. It is a desktop application for browsing the aligned or mapped reads [Ruffalo, M, LaFramboise, T, Koyutürk, M. Comparative analysis of algorithms for next-generation sequencing read alignment. Bioinformatics 2011;27:2790-6] at different levels of magnification, from nucleotide level, where the base qualities can be seen, to genome or chromosome level where overall coverage is shown. To enable in-depth investigation of NGS data, various views are provided that can be configured to highlight interesting aspects of the data. Multiple read alignment files can be overlaid to compare results from different experiments, and filters can be applied to facilitate the interpretation of the aligned reads. As well as being a standalone application it can be used as an integrated part of the Artemis genome browser, BamView allows the user to study NGS data in the context of the sequence and annotation of the reference genome. Single nucleotide polymorphism (SNP) density and candidate SNP sites can be highlighted and investigated, and read-pair information can be used to discover large structural insertions and deletions. The application will also calculate simple analyses of the read mapping, including reporting the read counts and reads per kilobase per million mapped reads (RPKM) for genes selected by the user. AVAILABILITY: BamView and Artemis are freely available software. These can be downloaded from their home pages: http://bamview.sourceforge.net/; http://www.sanger.ac.uk/resources/software/artemis/. Requirements: Java 1.6 or higher. Tim J. Carver, Simon R. Harris, Thomas D. Otto, Matthew Berriman, Julian Parkhill, Jacqueline A. Keane |
Briefings Bioinform. | 3 |
| 2010 | BamView: viewing mapped read alignment data in the context of the reference sequenceabstractSUMMARY: BamView is an interactive Java application for visualizing the large amounts of data stored for sequence reads which are aligned against a reference genome sequence. It supports the BAM (Binary Alignment/Map) format. It can be used in a number of contexts including SNP calling and structural annotation. BamView has also been integrated into Artemis so that the reads can be viewed in the context of the nucleotide sequence and genomic features. AVAILABILITY: BamView and Artemis are freely available (under a GPL licence) for download (for MacOSX, UNIX and Windows) at: http://bamview.sourceforge.net/ Tim J. Carver, Ulrike Böhme, Thomas D. Otto, Julian Parkhill, Matthew Berriman |
Bioinform. | 3 |
| 2010 | ProteinWorldDB: querying radical pairwise alignments among protein sets from complete genomesabstractMOTIVATION: Many analyses in modern biological research are based on comparisons between biological sequences, resulting in functional, evolutionary and structural inferences. When large numbers of sequences are compared, heuristics are often used resulting in a certain lack of accuracy. In order to improve and validate results of such comparisons, we have performed radical all-against-all comparisons of 4 million protein sequences belonging to the RefSeq database, using an implementation of the Smith-Waterman algorithm. This extremely intensive computational approach was made possible with the help of World Community Grid, through the Genome Comparison Project. The resulting database, ProteinWorldDB, which contains coordinates of pairwise protein alignments and their respective scores, is now made available. Users can download, compare and analyze the results, filtered by genomes, protein functions or clusters. ProteinWorldDB is integrated with annotations derived from Swiss-Prot, Pfam, KEGG, NCBI Taxonomy database and gene ontology. The database is a unique and valuable asset, representing a major effort to create a reliable and consistent dataset of cross-comparisons of the whole protein content encoded in hundreds of completely sequenced genomes using a rigorous dynamic programming approach. AVAILABILITY: The database can be accessed through http://proteinworlddb.org Thomas D. Otto, Marcos Catanho, Cristian Tristão, Márcia Bezerra, Renan Mathias Fernandes, Guilherme Steinberger Elias, Alexandre Capeletto Scaglia, Bill Bovermann, Viktors Berstis, Sérgio Lifschitz, Antonio B. de Miranda, Wim M. Degrave |
Bioinform. | 1 |
| 2010 | Iterative Correction of Reference Nucleotides (iCORN) using second generation sequencing technologyabstractMOTIVATION: The accuracy of reference genomes is important for downstream analysis but a low error rate requires expensive manual interrogation of the sequence. Here, we describe a novel algorithm (Iterative Correction of Reference Nucleotides) that iteratively aligns deep coverage of short sequencing reads to correct errors in reference genome sequences and evaluate their accuracy. RESULTS: Using Plasmodium falciparum (81% A + T content) as an extreme example, we show that the algorithm is highly accurate and corrects over 2000 errors in the reference sequence. We give examples of its application to numerous other eukaryotic and prokaryotic genomes and suggest additional applications. AVAILABILITY: The software is available at http://icorn.sourceforge.net Thomas D. Otto, Mandy Sanders, Matthew Berriman, Chris Newbold |
Bioinform. | 1 |
| 2009 | ABACAS: algorithm-based automatic contiguation of assembled sequencesabstractSUMMARY: Due to the availability of new sequencing technologies, we are now increasingly interested in sequencing closely related strains of existing finished genomes. Recently a number of de novo and mapping-based assemblers have been developed to produce high quality draft genomes from new sequencing technology reads. New tools are necessary to take contigs from a draft assembly through to a fully contiguated genome sequence. ABACAS is intended as a tool to rapidly contiguate (align, order, orientate), visualize and design primers to close gaps on shotgun assembled contigs based on a reference sequence. The input to ABACAS is a set of contigs which will be aligned to the reference genome, ordered and orientated, visualized in the ACT comparative browser, and optimal primer sequences are automatically generated. AVAILABILITY AND IMPLEMENTATION: ABACAS is implemented in Perl and is freely available for download from http://abacas.sourceforge.net. Samuel A. Assefa, Thomas M. Keane, Thomas D. Otto, Chris Newbold, Matthew Berriman |
Bioinform. | 3 |
| 2008 | ReRep: Computational detection of repetitive sequences in genome survey sequences (GSS)abstractBACKGROUND: Genome survey sequences (GSS) offer a preliminary global view of a genome since, unlike ESTs, they cover coding as well as non-coding DNA and include repetitive regions of the genome. A more precise estimation of the nature, quantity and variability of repetitive sequences very early in a genome sequencing project is of considerable importance, as such data strongly influence the estimation of genome coverage, library quality and progress in scaffold construction. Also, the elimination of repetitive sequences from the initial assembly process is important to avoid errors and unnecessary complexity. Repetitive sequences are also of interest in a variety of other studies, for instance as molecular markers. RESULTS: We designed and implemented a straightforward pipeline called ReRep, which combines bioinformatics tools for identifying repetitive structures in a GSS dataset. In a case study, we first applied the pipeline to a set of 970 GSSs, sequenced in our laboratory from the human pathogen Leishmania braziliensis, the causative agent of leishmaniosis, an important public health problem in Brazil. We also verified the applicability of ReRep to new sequencing technologies using a set of 454-reads of an Escheria coli. The behaviour of several parameters in the algorithm is evaluated and suggestions are made for tuning of the analysis. CONCLUSION: The ReRep approach for identification of repetitive elements in GSS datasets proved to be straightforward and efficient. Several potential repetitive sequences were found in a L. braziliensis GSS dataset generated in our laboratory, and further validated by the analysis of a more complete genomic dataset from the EMBL and Sanger Centre databases. ReRep also identified most of the E. coli K12 repeats prior to assembly in an example dataset obtained by automated sequencing using 454 technology. The parameters controlling the algorithm behaved consistently and may be tuned to the properties of the dataset, in particular to the length of sequencing reads and the genome coverage. ReRep is freely available for academic use at http://bioinfo.pdtis.fiocruz.br/ReRep/. Thomas D. Otto, Leonardo H. F. Gomes, Marcelo Alves-Ferreira, Antonio B. de Miranda, Wim M. Degrave |
BMC Bioinform. | 1 |
| 2008 | AnEnPi: identification and annotation of analogous enzymesabstractBACKGROUND: Enzymes are responsible for the catalysis of the biochemical reactions in metabolic pathways. Analogous enzymes are able to catalyze the same reactions, but they present no significant sequence similarity at the primary level, and possibly different tertiary structures as well. They are thought to have arisen as the result of independent evolutionary events. A detailed study of analogous enzymes may reveal new catalytic mechanisms, add information about the origin and evolution of biochemical pathways and disclose potential targets for drug development. RESULTS: In this work, we have constructed and implemented a new approach, AnEnPi (the Analogous Enzyme Pipeline), using a combination of bioinformatics tools like BLAST, HMMer, and in-house scripts, to assist in the identification, annotation, comparison and study of analogous and homologous enzymes. The algorithm for the detection of analogy is based i) on the construction of groups of homologous enzymes and ii) on the identification of cases where a given enzymatic activity is performed by two or more proteins without significant similarity between their primary structures. We applied this approach to a dataset obtained from KEGG Comprising all annotated enzymes, which resulted in the identification of 986 EC classes where putative analogy was detected (40.5% of all EC classes). AnEnPi is of considerable value in the construction of initial datasets that can be further curated, particularly in gene and genome annotation, in studies involving molecular evolution and metabolism and in the identification of new potential drug targets. CONCLUSION: AnEnPi is an efficient tool for detection and annotation of analogous enzymes and other enzymes in whole genomes. It is available for academic use at: http://bioinfo.pdtis.fiocruz.br/AnEnPi/ Thomas D. Otto, Ana Carolina R. Guimarães, Wim M. Degrave, Antonio B. de Miranda |
BMC Bioinform. | 1 |
| 2003 | Model-Free Functional MRI Analysis Using Topographic Independent Component Analysis
Anke Meyer-Bäse, Thomas D. Otto, Thomas Martinetz, Dorothee Auer, Axel Wismüller |
ESANN | 2 |