EDBT 2026 Demo / reviewers in the wild / expert
Travis J. Wheeler
dblp:74/3170 · also Travis John Wheeler
· DBLP profile ↗
10ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-2004-1785ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 100% | |
| Artificial intelligence
1 paper |
Representation and self-supervised learning · 50% Language models and text generation · 50% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% |
Topics — the 14 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › protein sequence analysis
protein homology detection |
0.9 | 1 | 2025 | NEAR: neural embeddings for amino acid relationships · Bioinform. 2025 |
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model |
0.9 | 1 | 2025 | NEAR: neural embeddings for amino acid relationships · Bioinform. 2025 |
Machine learning › Representation and self-supervised learning › latent space
latent space analysis |
0.7 | 1 | 2023 | Reliable Measures of Spread in High Dimensional Latent Spaces · ICML 2023 |
Storage systems
data compression |
0.3 | 1 | 2026 | MDCompress: better, faster compression of molecular dynamics simulation trajectories · Bioinform. 2026 |
Bioinformatics and computational biology › sequence analysis › sequence similarity search
DNA similarity search |
0.2 | 1 | 2013 | nhmmer: DNA homology search with profile HMMs · Bioinform. 2013 |
Bioinformatics and computational biology › sequence analysis › profile hidden markov model
profile hidden markov model search |
0.2 | 1 | 2013 | nhmmer: DNA homology search with profile HMMs · Bioinform. 2013 |
Bioinformatics and computational biology
sequence analysis |
0.2 | 1 | 2013 | nhmmer: DNA homology search with profile HMMs · Bioinform. 2013 |
Bioinformatics and computational biology › sequence analysis
sequence similarity search |
0.2 | 1 | 2013 | nhmmer: DNA homology search with profile HMMs · Bioinform. 2013 |
Bioinformatics and computational biology › multiple sequence alignment
alignment quality assessment |
0.1 | 1 | 2012 | Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising · RECOMB 2012 |
Bioinformatics and computational biology
multiple sequence alignment |
0.1 | 1 | 2012 | Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising · RECOMB 2012 |
Bioinformatics and computational biology
sequence alignment |
0.1 | 1 | 2012 | Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising · RECOMB 2012 |
Bioinformatics and computational biology
protein structure prediction |
0.1 | 1 | 2009 | Learning Models for Aligning Protein Sequences with Predicted Secondary Structure · RECOMB 2009 |
Bioinformatics and computational biology › structural bioinformatics
sequence-structure alignment |
0.1 | 1 | 2009 | Learning Models for Aligning Protein Sequences with Predicted Secondary Structure · RECOMB 2009 |
Bioinformatics and computational biology › sequence analysis › sequencing data processing
sequence quality control |
0.1 | 1 | 2005 | Evaluating and improving cDNA sequence quality with cQC · Bioinform. 2005 |
Methods — techniques the papers use, named apart from their topics
random-access decompression · 2.0multithreading · 2.0resnet · 0.9k-NN search · 0.9contrastive learning · 0.9principal component analysis · 0.7entropy-based measure · 0.7profile hidden markov model · 0.2probabilistic inference · 0.2secondary structure prediction · 0.1machine learning · 0.1sequence alignment · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MDCompress: better, faster compression of molecular dynamics simulation trajectoriesabstractMOTIVATION: Molecular dynamics (MD) simulations model the physical movements of atoms in biomolecular systems over time, providing atomic-resolution insight into conformational changes, binding events, and dynamic behaviors that cannot be captured by static structures alone. As such, MD simulations are playing an increasingly important role in understanding the functional roles and molecular interactions of proteins. However, trajectories from these simulations can be extremely large, often reaching tens of gigabytes for a single simulation of modest duration. This creates substantial challenges for storage and data transfer, motivating efficient compression strategies. Furthermore, many downstream analyses require extraction of only a subset of frames or specific atoms from the full trajectory, so an ideal compression format should support rapid random-access decompression of such samplings without requiring full file decompression. RESULTS: Here, we introduce MDCompress, a new trajectory compression format and accompanying software implementation that meets these goals. MDCompress produces compressed trajectory files that are 15-37% smaller than those generated by the widely-used XTC format, while achieving faster compression and decompression speeds through efficient multithreading. AVAILABILITY AND IMPLEMENTATION: The MDCompress software and library are released under an open license (BSD-3) and may be downloaded at https://github.com/refresh-bio/mdcompress and is also available as a Zenodo repository at 10.5281/zenodo.19218347. Marek Kokot, Amitava Roy, Travis J. Wheeler, Sebastian Deorowicz |
Bioinform. | 3 |
| 2025 | NEAR: neural embeddings for amino acid relationshipsabstractSUMMARY: Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present Neural Embeddings for Amino acid Relationships (NEAR), a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR's ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript, we focus on a more straightforward analysis of the model's value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail. AVAILABILITY AND IMPLEMENTATION: NEAR is under an open-source license. Code and data curation instructions can be found at https://github.com/TravisWheelerLab/NEAR. Daniel Olson, Thomas Colligan, Daphne Demekas, Jack W. Roddy, Ken Youens-Clark, Travis J. Wheeler |
Bioinform. | 6 |
| 2024 | An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov modelsabstractAbstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline. Travis J. Wheeler |
BMC Bioinform. | 2 |
| 2023 | Reliable Measures of Spread in High Dimensional Latent SpacesabstractUnderstanding geometric properties of the latent spaces of natural language processing models allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model’s latent space, or how fully the available latent space is being used. We demonstrate that the commonly used measures of data spread, average cosine similarity and a partition function min/max ratio I(V), do not provide reliable metrics to compare the use of latent space across data distributions. We propose and examine six alternative measures of data spread, all of which improve over these current metrics when applied to seven synthetic data distributions. Of our proposed measures, we recommend one principal component-based measure and one entropy-based measure that provide reliable, relative measures of spread and can be used to compare models of different sizes and dimensionalities. Anna C. Marbut, Katy McKinney-Bock, Travis J. Wheeler |
ICML | 3 |
| 2014 | Skylign: a tool for creating informative, interactive logos representing sequence alignments and profile hidden Markov modelsabstractBACKGROUND: Logos are commonly used in molecular biology to provide a compact graphical representation of the conservation pattern of a set of sequences. They render the information contained in sequence alignments or profile hidden Markov models by drawing a stack of letters for each position, where the height of the stack corresponds to the conservation at that position, and the height of each letter within a stack depends on the frequency of that letter at that position. RESULTS: We present a new tool and web server, called Skylign, which provides a unified framework for creating logos for both sequence alignments and profile hidden Markov models. In addition to static image files, Skylign creates a novel interactive logo plot for inclusion in web pages. These interactive logos enable scrolling, zooming, and inspection of underlying values. Skylign can avoid sampling bias in sequence alignments by down-weighting redundant sequences and by combining observed counts with informed priors. It also simplifies the representation of gap parameters, and can optionally scale letter heights based on alternate calculations of the conservation of a position. CONCLUSION: Skylign is available as a website, a scriptable web service with a RESTful interface, and as a software package for download. Skylign's interactive logos are easily incorporated into a web page with just a few lines of HTML markup. Skylign may be found at http://skylign.org. Travis J. Wheeler, Jody Clements, Robert D. Finn |
BMC Bioinform. | 1 |
| 2013 | nhmmer: DNA homology search with profile HMMsabstractSUMMARY: Sequence database searches are an essential part of molecular biology, providing information about the function and evolutionary history of proteins, RNA molecules and DNA sequence elements. We present a tool for DNA/DNA sequence comparison that is built on the HMMER framework, which applies probabilistic inference methods based on hidden Markov models to the problem of homology search. This tool, called nhmmer, enables improved detection of remote DNA homologs, and has been used in combination with Dfam and RepeatMasker to improve annotation of transposable elements in the human genome. AVAILABILITY: nhmmer is a part of the new HMMER3.1 release. Source code and documentation can be downloaded from http://hmmer.org. HMMER3.1 is freely licensed under the GNU GPLv3 and should be portable to any POSIX-compliant operating system, including Linux and Mac OS/X. Travis J. Wheeler, Sean R. Eddy |
Bioinform. | 1 |
| 2012 | Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising
Dan F. DeBlasio, Travis J. Wheeler, John D. Kececioglu |
RECOMB | 2 |
| 2009 | Learning Models for Aligning Protein Sequences with Predicted Secondary Structure
Eagu Kim, Travis J. Wheeler, John D. Kececioglu |
RECOMB | 2 |
| 2009 | Large-Scale Neighbor-Joining with NINJA
Travis J. Wheeler |
WABI | 1 |
| 2005 | Evaluating and improving cDNA sequence quality with cQCabstractAbstract Summary: Errors are prevalent in cDNA sequences but the extent to which sequence collections differ in frequencies and types of errors has not been investigated systematically. cDNA quality control, or cQC, was developed to evaluate the quality of cDNA sequence collections and to revise those sequences that differ from a higher quality genomic sequence. After removing rRNA, vector, bacterial insertion sequence and chimeric cDNA contaminants, small-scale nucleotide discrepancies were found in 51% of cDNA sequences from one Arabidopsis cDNA collection, 89% from a second Arabidopsis collection and 75% from a rice collection. These errors created premature termination codons in 4 and 42% of cDNA sequences in the respective Arabidopsis collections and in 7% of the rice cDNA sequences. Availability: A web-based version of cQC, source code and revised cDNA collections are available at Contact: [email protected] Supplementary information: Further text, tables and figures are available at the above website or on Bioinformatics online. Celine A. Hayden, Travis J. Wheeler, Richard A. Jorgensen |
Bioinform. | 2 |