Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Travis J. Wheeler

dblp:74/3170 · also Travis John Wheeler · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
4since 2021 · last 2026
0000-0003-2004-1785ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
6 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Representation and self-supervised learning · 50% Language models and text generation · 50%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Storage systems · 100%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology › protein sequence analysis
protein homology detection
0.912025
NEAR: neural embeddings for amino acid relationships · Bioinform. 2025
Bioinformatics and computational biology › protein sequence analysis › protein sequence representation
protein language model
0.912025
NEAR: neural embeddings for amino acid relationships · Bioinform. 2025
Machine learning › Representation and self-supervised learning › latent space
latent space analysis
0.712023
Reliable Measures of Spread in High Dimensional Latent Spaces · ICML 2023
Storage systems
data compression
0.312026
MDCompress: better, faster compression of molecular dynamics simulation trajectories · Bioinform. 2026
Bioinformatics and computational biology › sequence analysis › sequence similarity search
DNA similarity search
0.212013
nhmmer: DNA homology search with profile HMMs · Bioinform. 2013
Bioinformatics and computational biology › sequence analysis › profile hidden markov model
profile hidden markov model search
0.212013
nhmmer: DNA homology search with profile HMMs · Bioinform. 2013
Bioinformatics and computational biology
sequence analysis
0.212013
nhmmer: DNA homology search with profile HMMs · Bioinform. 2013
Bioinformatics and computational biology › sequence analysis
sequence similarity search
0.212013
nhmmer: DNA homology search with profile HMMs · Bioinform. 2013
Bioinformatics and computational biology › multiple sequence alignment
alignment quality assessment
0.112012
Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising · RECOMB 2012
Bioinformatics and computational biology
multiple sequence alignment
0.112012
Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising · RECOMB 2012
Bioinformatics and computational biology
sequence alignment
0.112012
Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising · RECOMB 2012
Bioinformatics and computational biology
protein structure prediction
0.112009
Learning Models for Aligning Protein Sequences with Predicted Secondary Structure · RECOMB 2009
Bioinformatics and computational biology › structural bioinformatics
sequence-structure alignment
0.112009
Learning Models for Aligning Protein Sequences with Predicted Secondary Structure · RECOMB 2009
Bioinformatics and computational biology › sequence analysis › sequencing data processing
sequence quality control
0.112005
Evaluating and improving cDNA sequence quality with cQC · Bioinform. 2005

Methods — techniques the papers use, named apart from their topics

random-access decompression · 2.0multithreading · 2.0resnet · 0.9k-NN search · 0.9contrastive learning · 0.9principal component analysis · 0.7entropy-based measure · 0.7profile hidden markov model · 0.2probabilistic inference · 0.2secondary structure prediction · 0.1machine learning · 0.1sequence alignment · 0.1
YearPublicationVenuePosition
2026 MDCompress: better, faster compression of molecular dynamics simulation trajectories
abstract
MOTIVATION: Molecular dynamics (MD) simulations model the physical movements of atoms in biomolecular systems over time, providing atomic-resolution insight into conformational changes, binding events, and dynamic behaviors that cannot be captured by static structures alone. As such, MD simulations are playing an increasingly important role in understanding the functional roles and molecular interactions of proteins. However, trajectories from these simulations can be extremely large, often reaching tens of gigabytes for a single simulation of modest duration. This creates substantial challenges for storage and data transfer, motivating efficient compression strategies. Furthermore, many downstream analyses require extraction of only a subset of frames or specific atoms from the full trajectory, so an ideal compression format should support rapid random-access decompression of such samplings without requiring full file decompression. RESULTS: Here, we introduce MDCompress, a new trajectory compression format and accompanying software implementation that meets these goals. MDCompress produces compressed trajectory files that are 15-37% smaller than those generated by the widely-used XTC format, while achieving faster compression and decompression speeds through efficient multithreading. AVAILABILITY AND IMPLEMENTATION: The MDCompress software and library are released under an open license (BSD-3) and may be downloaded at https://github.com/refresh-bio/mdcompress and is also available as a Zenodo repository at 10.5281/zenodo.19218347.
Marek Kokot, Amitava Roy, Travis J. Wheeler, Sebastian Deorowicz
Bioinform.3
2025 NEAR: neural embeddings for amino acid relationships
abstract
SUMMARY: Protein language models (PLMs) have recently demonstrated potential to supplant classical protein database search methods based on sequence alignment, but are slower than common alignment-based tools and appear to be prone to a high rate of false labeling. Here, we present Neural Embeddings for Amino acid Relationships (NEAR), a method based on neural representation learning that is designed to improve both speed and accuracy of search for likely homologs in a large protein sequence database. NEAR's ResNet embedding model is trained using contrastive learning guided by trusted sequence alignments. It computes per-residue embeddings for target and query protein sequences, and identifies alignment candidates with a pipeline consisting of residue-level k-NN search and a simple neighbor aggregation scheme. Tests on a benchmark consisting of trusted remote homologs and randomly shuffled decoy sequences reveal that NEAR substantially improves accuracy relative to state-of-the-art PLMs, with lower memory requirements and faster embedding and search speed. While these results suggest that the NEAR model may be useful for standalone homology detection with increased sensitivity over standard alignment-based methods, in this manuscript, we focus on a more straightforward analysis of the model's value as a high-speed pre-filter for sensitive annotation. In that context, NEAR is at least 5x faster than the pre-filter currently used in the widely used profile hidden Markov model (pHMM) search tool HMMER3, and also outperforms the pre-filter used in our fast pHMM tool, nail. AVAILABILITY AND IMPLEMENTATION: NEAR is under an open-source license. Code and data curation instructions can be found at https://github.com/TravisWheelerLab/NEAR.
Daniel Olson, Thomas Colligan, Daphne Demekas, Jack W. Roddy, Ken Youens-Clark, Travis J. Wheeler
Bioinform.6
2024 An FPGA-based hardware accelerator supporting sensitive sequence homology filtering with profile hidden Markov models
abstract
Abstract Background Sequence alignment lies at the heart of genome sequence annotation. While the BLAST suite of alignment tools has long held an important role in alignment-based sequence database search, greater sensitivity is achieved through the use of profile hidden Markov models (pHMMs). Here, we describe an FPGA hardware accelerator, called HAVAC, that targets a key bottleneck step (SSV) in the analysis pipeline of the popular pHMM alignment tool, HMMER. Results The HAVAC kernel calculates the SSV matrix at 1739 GCUPS on a $$\sim$$ ∼ $3000 Xilinx Alveo U50 FPGA accelerator card, $$\sim$$ ∼ 227× faster than the optimized SSV implementation in nhmmer . Accounting for PCI-e data transfer data processing, HAVAC is 65× faster than nhmmer’s SSV with one thread and 35× faster than nhmmer with four threads, and uses $$\sim$$ ∼ 31% the energy of a traditional high end Intel CPU. Conclusions HAVAC demonstrates the potential offered by FPGA hardware accelerators to produce dramatic speed gains in sequence annotation and related bioinformatics applications. Because these computations are performed on a co-processor, the host CPU remains free to simultaneously compute other aspects of the analysis pipeline.
Travis J. Wheeler
BMC Bioinform.2
2023 Reliable Measures of Spread in High Dimensional Latent Spaces
abstract
Understanding geometric properties of the latent spaces of natural language processing models allows the manipulation of these properties for improved performance on downstream tasks. One such property is the amount of data spread in a model’s latent space, or how fully the available latent space is being used. We demonstrate that the commonly used measures of data spread, average cosine similarity and a partition function min/max ratio I(V), do not provide reliable metrics to compare the use of latent space across data distributions. We propose and examine six alternative measures of data spread, all of which improve over these current metrics when applied to seven synthetic data distributions. Of our proposed measures, we recommend one principal component-based measure and one entropy-based measure that provide reliable, relative measures of spread and can be used to compare models of different sizes and dimensionalities.
Anna C. Marbut, Katy McKinney-Bock, Travis J. Wheeler
ICML3
2014 Skylign: a tool for creating informative, interactive logos representing sequence alignments and profile hidden Markov models
abstract
BACKGROUND: Logos are commonly used in molecular biology to provide a compact graphical representation of the conservation pattern of a set of sequences. They render the information contained in sequence alignments or profile hidden Markov models by drawing a stack of letters for each position, where the height of the stack corresponds to the conservation at that position, and the height of each letter within a stack depends on the frequency of that letter at that position. RESULTS: We present a new tool and web server, called Skylign, which provides a unified framework for creating logos for both sequence alignments and profile hidden Markov models. In addition to static image files, Skylign creates a novel interactive logo plot for inclusion in web pages. These interactive logos enable scrolling, zooming, and inspection of underlying values. Skylign can avoid sampling bias in sequence alignments by down-weighting redundant sequences and by combining observed counts with informed priors. It also simplifies the representation of gap parameters, and can optionally scale letter heights based on alternate calculations of the conservation of a position. CONCLUSION: Skylign is available as a website, a scriptable web service with a RESTful interface, and as a software package for download. Skylign's interactive logos are easily incorporated into a web page with just a few lines of HTML markup. Skylign may be found at http://skylign.org.
Travis J. Wheeler, Jody Clements, Robert D. Finn
BMC Bioinform.1
2013 nhmmer: DNA homology search with profile HMMs
abstract
SUMMARY: Sequence database searches are an essential part of molecular biology, providing information about the function and evolutionary history of proteins, RNA molecules and DNA sequence elements. We present a tool for DNA/DNA sequence comparison that is built on the HMMER framework, which applies probabilistic inference methods based on hidden Markov models to the problem of homology search. This tool, called nhmmer, enables improved detection of remote DNA homologs, and has been used in combination with Dfam and RepeatMasker to improve annotation of transposable elements in the human genome. AVAILABILITY: nhmmer is a part of the new HMMER3.1 release. Source code and documentation can be downloaded from http://hmmer.org. HMMER3.1 is freely licensed under the GNU GPLv3 and should be portable to any POSIX-compliant operating system, including Linux and Mac OS/X.
Travis J. Wheeler, Sean R. Eddy
Bioinform.1
2012 Estimating the Accuracy of Multiple Alignments and its Use in Parameter Advising
Dan F. DeBlasio, Travis J. Wheeler, John D. Kececioglu
RECOMB2
2009 Learning Models for Aligning Protein Sequences with Predicted Secondary Structure
Eagu Kim, Travis J. Wheeler, John D. Kececioglu
RECOMB2
2009 Large-Scale Neighbor-Joining with NINJA
Travis J. Wheeler
WABI1
2005 Evaluating and improving cDNA sequence quality with cQC
abstract
Abstract Summary: Errors are prevalent in cDNA sequences but the extent to which sequence collections differ in frequencies and types of errors has not been investigated systematically. cDNA quality control, or cQC, was developed to evaluate the quality of cDNA sequence collections and to revise those sequences that differ from a higher quality genomic sequence. After removing rRNA, vector, bacterial insertion sequence and chimeric cDNA contaminants, small-scale nucleotide discrepancies were found in 51% of cDNA sequences from one Arabidopsis cDNA collection, 89% from a second Arabidopsis collection and 75% from a rice collection. These errors created premature termination codons in 4 and 42% of cDNA sequences in the respective Arabidopsis collections and in 7% of the rice cDNA sequences. Availability: A web-based version of cQC, source code and revised cDNA collections are available at Contact: [email protected] Supplementary information: Further text, tables and figures are available at the above website or on Bioinformatics online.
Celine A. Hayden, Travis J. Wheeler, Richard A. Jorgensen
Bioinform.2