Martin Steinegger

dblp:179/5577 · DBLP profile ↗
← Back
12ranked-venue papers
1as first author
7since 2021 · last 2026
0000-0001-8781-9753ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 11 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Accelign: a GPU-based library for accelerating pairwise sequence alignment
abstract
BACKGROUND: The continually increasing volume of sequence data results in a growing demand for fast implementations of core algorithms. Computation of pairwise alignments based on dynamic programming is an important part in many bioinformatics pipelines and a major contributor to overall runtime due to the associated quadratic time complexity. This motivates the need for a library of efficient implementations on modern GPUs for a variety of alignment algorithms for different types of sequence data including DNA, RNA, and proteins. RESULTS: Accelign is a library of accelerated pairwise sequence alignment algorithms for CUDA-enabled GPUs. Its parallelization strategy is based on a common wavefront design that can be adapted to support a variety of dynamic programming algorithms: local, global, and semi-global alignment of genomic and protein sequences with a variety of commonly used scoring schemes supporting one-to-one, one-to-many or all-to-all pairwise sequence alignments. This leads to a peak performance between 16.1 TCUPS and 9.1 TCUPS for computing optimal global alignment scores with linear gaps and affine gap penalties on a single RTX PRO 6000 Blackwell GPU, respectively. In addition, our library demonstrates significant speedups in several real-world case studies over prior CPU-based (SeqAn, Parasail, BSalign, EdLib, KSW2, WFA2, A*PA2) and GPU-based libraries (ADEPT, GASAL2), and can even outperform highly customized algorithms (WFA-GPU, CUDASW++4.0). Furthermore, the performance of our approach scales linearly with the number of employed GPUs, which makes it feasible to exploit multi-GPU nodes for increased processing speeds. CONCLUSION: Accelign provides significant speedups for commonly used pairwise alignment algorithms compared to prior implementations. It is freely available at https://github.com/fkallen/Accelign .
Felix Kallenborn, Fawaz Dabbaghie, Martin Steinegger, Bertil Schmidt
BMC Bioinform.3
2025 Easy and interactive taxonomic profiling with Metabuli App
abstract
SUMMARY: Accurate metagenomic taxonomic profiling is critical for understanding microbial communities. However, computational analysis often requires command-line proficiency and high-performance computing resources. To lower these barriers, we developed Metabuli App, an all-in-one desktop application that efficiently runs taxonomic profiling locally on a consumer-grade computer. It features user-friendly graphical interfaces for custom database curation, raw read quality control (QC), taxonomic profiling, and interactive result visualization. AVAILABILITY AND IMPLEMENTATION: GPLv3-licensed source code and prebuilt apps for Windows, macOS, and Linux are available at https://github.com/steineggerlab/Metabuli-App and are archived at https://doi.org/10.5281/zenodo.15876171. Analysis scripts are available at https://github.com/jaebeom-kim/metabuli-app-analysis. The Sankey-based taxonomy visualization component is available at https://github.com/steineggerlab/taxoview for easy integration into other web projects.
Sunjae Lee, Jaebeom Kim, Milot Mirdita, Cameron L. M. Gilchrist, Martin Steinegger
Bioinform.5
2023 Foldcomp: a library and format for compressing and indexing large protein structure sets
abstract
SUMMARY: Highly accurate protein structure predictors have generated hundreds of millions of protein structures; these pose a challenge in terms of storage and processing. Here, we present Foldcomp, a novel lossy structure compression algorithm, and indexing system to address this challenge. By using a combination of internal and Cartesian coordinates and a bi-directional NeRF-based strategy, Foldcomp improves the compression ratio by a factor of three compared to the next best method. Its reconstruction error of 0.08 Å is comparable to the best lossy compressor. It is five times faster than the next fastest compressor and competes with the fastest decompressors. With its multi-threading implementation and a Python interface that allows for easy database downloads and efficient querying of protein structures by accession, Foldcomp is a powerful tool for managing and analysing large collections of protein structures. AVAILABILITY AND IMPLEMENTATION: Foldcomp is a free open-source software (GPLv3) and available for Linux, macOS, and Windows at https://foldcomp.foldseek.com. Foldcomp provides the AlphaFold Swiss-Prot (2.9GB), TrEMBL (1.1TB), and ESMatlas HQ (114GB) database ready-for-download.
Hyunbin Kim, Milot Mirdita, Martin Steinegger
Bioinform.3
2023 Block Aligner: an adaptive SIMD-accelerated aligner for sequences and position-specific scoring matrices
abstract
MOTIVATION: Efficiently aligning sequences is a fundamental problem in bioinformatics. Many recent algorithms for computing alignments through Smith-Waterman-Gotoh dynamic programming (DP) exploit Single Instruction Multiple Data (SIMD) operations on modern CPUs for speed. However, these advances have largely ignored difficulties associated with efficiently handling complex scoring matrices or large gaps (insertions or deletions). RESULTS: We propose a new SIMD-accelerated algorithm called Block Aligner for aligning nucleotide and protein sequences against other sequences or position-specific scoring matrices. We introduce a new paradigm that uses blocks in the DP matrix that greedily shift, grow, and shrink. This approach allows regions of the DP matrix to be adaptively computed. Our algorithm reaches over 5-10 times faster than some previous methods while incurring an error rate of less than 3% on protein and long read datasets, despite large gaps and low sequence identities. AVAILABILITY AND IMPLEMENTATION: Our algorithm is implemented for global, local, and X-drop alignments. It is available as a Rust library (with C bindings) at https://github.com/Daniel-Liu-c0deb0t/block-aligner.
Daniel Liu, Martin Steinegger
Bioinform.2
2022 PhyloCSF++: a fast and user-friendly implementation of PhyloCSF with annotation tools
abstract
SUMMARY: PhyloCSF++ is an efficient and parallelized C++ implementation of the popular PhyloCSF method to distinguish protein-coding and non-coding regions in a genome based on multiple sequence alignments (MSAs). It can score alignments or produce browser tracks for entire genomes in the wig file format. Additionally, PhyloCSF++ annotates coding sequences in GFF/GTF files using precomputed tracks or computes and scores MSAs on the fly with MMseqs2. AVAILABILITY AND IMPLEMENTATION: PhyloCSF++ is released under the AGPLv3 license. Binaries and source code are available at https://github.com/cpockrandt/PhyloCSFpp. The software can be installed through bioconda. A variety of tracks can be accessed through ftp://ftp.ccb.jhu.edu/pub/software/phylocsfpp/.
Christopher Pockrandt, Martin Steinegger, Steven Salzberg
Bioinform.2
2022 ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
abstract
Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The protein LMs (pLMs) were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw pLM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks: (1) a per-residue (per-token) prediction of protein secondary structure (3-state accuracy Q3=81%-87%); (2) per-protein (pooling) predictions of protein sub-cellular location (ten-state accuracy: Q10=81%) and membrane versus water-soluble (2-state accuracy Q2=91%). For secondary structure, the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without multiple sequence alignments (MSAs) or evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that pLMs learned some of the grammar of the language of life. All our models are available through https://github.com/agemagician/ProtTrans.
Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang 0008, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, Burkhard Rost
IEEE Trans. Pattern Anal. Mach. Intell.10
2021 Fast and sensitive taxonomic assignment to metagenomic contigs
abstract
SUMMARY: MMseqs2 taxonomy is a new tool to assign taxonomic labels to metagenomic contigs. It extracts all possible protein fragments from each contig, quickly retains those that can contribute to taxonomic annotation, assigns them with robust labels and determines the contig's taxonomic identity by weighted voting. Its fragment extraction step is suitable for the analysis of all domains of life. MMseqs2 taxonomy is 2-18× faster than state-of-the-art tools and also contains new modules for creating and manipulating taxonomic reference databases as well as reporting and visualizing taxonomic assignments. AVAILABILITY AND IMPLEMENTATION: MMseqs2 taxonomy is part of the MMseqs2 free open-source software package available for Linux, macOS and Windows at https://mmseqs.com. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Milot Mirdita, Martin Steinegger, Florian P. Breitwieser, Johannes Söding, Eli Levy Karin
Bioinform.2
2019 MMseqs2 desktop and local web server app for fast, interactive sequence searches
abstract
SUMMARY: The MMseqs2 desktop and web server app facilitates interactive sequence searches through custom protein sequence and profile databases on personal workstations. By eliminating MMseqs2's runtime overhead, we reduced response times to a few seconds at sensitivities close to BLAST. AVAILABILITY AND IMPLEMENTATION: The app is easy to install for non-experts. GPLv3-licensed code, pre-built desktop app packages for Windows, MacOS and Linux, Docker images for the web server application and a demo web server are available at https://search.mmseqs.com. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Milot Mirdita, Martin Steinegger, Johannes Söding
Bioinform.2
2019 HH-suite3 for fast remote homology detection and deep protein annotation
abstract
BACKGROUND: HH-suite is a widely used open source software suite for sensitive sequence similarity searches and protein fold recognition. It is based on pairwise alignment of profile Hidden Markov models (HMMs), which represent multiple sequence alignments of homologous proteins. RESULTS: We developed a single-instruction multiple-data (SIMD) vectorized implementation of the Viterbi algorithm for profile HMM alignment and introduced various other speed-ups. These accelerated the search methods HHsearch by a factor 4 and HHblits by a factor 2 over the previous version 2.0.16. HHblits3 is ∼10× faster than PSI-BLAST and ∼20× faster than HMMER3. Jobs to perform HHsearch and HHblits searches with many query profile HMMs can be parallelized over cores and over cluster servers using OpenMP and message passing interface (MPI). The free, open-source, GPLv3-licensed software is available at https://github.com/soedinglab/hh-suite . CONCLUSION: The added functionalities and increased speed of HHsearch and HHblits should facilitate their use in large-scale protein structure and function prediction, e.g. in metagenomics and genomics projects.
Martin Steinegger, Markus Meier, Milot Mirdita, Harald Vöhringer, Stefan J. Haunsberger, Johannes Söding
BMC Bioinform.1
2018 HFSP: high speed homology-driven function annotation of proteins
abstract
Motivation: The rapid drop in sequencing costs has produced many more (predicted) protein sequences than can feasibly be functionally annotated with wet-lab experiments. Thus, many computational methods have been developed for this purpose. Most of these methods employ homology-based inference, approximated via sequence alignments, to transfer functional annotations between proteins. The increase in the number of available sequences, however, has drastically increased the search space, thus significantly slowing down alignment methods. Results: Here we describe homology-derived functional similarity of proteins (HFSP), a novel computational method that uses results of a high-speed alignment algorithm, MMseqs2, to infer functional similarity of proteins on the basis of their alignment length and sequence identity. We show that our method is accurate (85% precision) and fast (more than 40-fold speed increase over state-of-the-art). HFSP can help correct at least a 16% error in legacy curations, even for a resource of as high quality as Swiss-Prot. These findings suggest HFSP as an ideal resource for large-scale functional annotation efforts. Supplementary information: Supplementary data are available at Bioinformatics online.
Yannick Mahlich, Martin Steinegger, Burkhard Rost, Yana Bromberg
Bioinform.2
2017 Updating lidar-derived crown cover density products with sentinel-2
abstract
Crown cover density (CCD) is one important forest attribute used in forest management. With remote sensing, crown cover density maps can be derived from the spectral information of optical satellite imagery or from a normalized digital surface model (nDSM). LiDAR data based applications provide the most accurate results, but LiDAR campaigns are expensive and available data are often outdated. We propose a new method to update LiDAR-derived CCD products and to map forest change. The method is based on Sentinel-2 imagery and an outdated LiDAR nDSM used to train a kNN classifier. CCD estimations are derived for two tests sites in Austria. Results are compared with the LiDAR CCD values in unchanged forest and with the latest tree cover density product of the European Copernicus High Resolution Layers Forest. Results demonstrate the operability of the workflow. User accuracies for forest change detection are very high with 87.3% and 94.8%.
Janik Deutscher, Klaus Granica, Martin Steinegger, Manuela Hirschmugl, Roland Perko, Mathias Schardt
IGARSS3
2016 MMseqs software suite for fast and deep clustering and searching of large protein sequence sets
abstract
MOTIVATION: Sequence databases are growing fast, challenging existing analysis pipelines. Reducing the redundancy of sequence databases by similarity clustering improves speed and sensitivity of iterative searches. But existing tools cannot efficiently cluster databases of the size of UniProt to 50% maximum pairwise sequence identity or below. Furthermore, in metagenomics experiments typically large fractions of reads cannot be matched to any known sequence anymore because searching with sensitive but relatively slow tools (e.g. BLAST or HMMER3) through comprehensive databases such as UniProt is becoming too costly. RESULTS: MMseqs (Many-against-Many sequence searching) is a software suite for fast and deep clustering and searching of large datasets, such as UniProt, or 6-frame translated metagenomics sequencing reads. MMseqs contains three core modules: a fast and sensitive prefiltering module that sums up the scores of similar k-mers between query and target sequences, an SSE2- and multi-core-parallelized local alignment module, and a clustering module.In our homology detection benchmarks, MMseqs is much more sensitive and 4-30 times faster than UBLAST and RAPsearch, respectively, although it does not reach BLAST sensitivity yet. Using its cascaded clustering workflow, MMseqs can cluster large databases down to ∼30% sequence identity at hundreds of times the speed of BLASTclust and much deeper than CD-HIT and USEARCH. MMseqs can also update a database clustering in linear instead of quadratic time. Its much improved sensitivity-speed trade-off should make MMseqs attractive for a wide range of large-scale sequence analysis tasks. AVAILABILITY AND IMPLEMENTATION: MMseqs is open-source software available under GPL at https://github.com/soedinglab/MMseqs CONTACT: [email protected], [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Maria Hauser, Martin Steinegger, Johannes Söding
Bioinform.2