Kai Dührkop

dblp:56/11538 · DBLP profile ↗
← Back
10ranked-venue papers
4as first author
2since 2021 · last 2024
0000-0002-9056-0540ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 8 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Theory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Interdisciplinary, comprehensive, and emerging computing
7 papers
Bioinformatics and computational biology · 100%
Artificial intelligence
1 paper
Deep learning architectures and training · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Bioinformatics and computational biology
metabolomics
2.262024
MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024
Deep kernel learning improves molecular fingerprint prediction from tandem mass spectra · Bioinform. 2022
Bayesian networks for mass spectrometric metabolite identification via molecular fingerprints · Bioinform. 2018
Bioinformatics and computational biology › proteomics
mass spectrometry data analysis
1.022024
MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024
Fragmentation Trees Reloaded · RECOMB 2015
Bioinformatics and computational biology › metabolomics
molecular fingerprint prediction
0.832022
Deep kernel learning improves molecular fingerprint prediction from tandem mass spectra · Bioinform. 2022
Metabolite identification through multiple kernel learning on fragmentation trees · Bioinform. 2014
Fast metabolite identification with Input Output Kernel Regression · Bioinform. 2016
Bioinformatics and computational biology › metabolomics
metabolite identification
0.832018
Bayesian networks for mass spectrometric metabolite identification via molecular fingerprints · Bioinform. 2018
Fast metabolite identification with Input Output Kernel Regression · Bioinform. 2016
Metabolite identification through multiple kernel learning on fragmentation trees · Bioinform. 2014
Bioinformatics and computational biology › structural biology
molecular structure determination
0.812024
MassSpecGym: A benchmark for the discovery and identification of molecules · NeurIPS 2024
Bioinformatics and computational biology
proteomics
0.212015
Fragmentation Trees Reloaded · RECOMB 2015
Bioinformatics and computational biology › proteomics
tandem mass spectrometry
0.112018
Bayesian networks for mass spectrometric metabolite identification via molecular fingerprints · Bioinform. 2018
Bioinformatics and computational biology › structural bioinformatics
molecular structure prediction
0.112016
Fast metabolite identification with Input Output Kernel Regression · Bioinform. 2016

Methods — techniques the papers use, named apart from their topics

spectrum simulation · 1.5molecule retrieval · 1.5de novo molecular structure generation · 1.5nyström approximation · 0.6kernel support vector machine · 0.6deep neural network · 0.6machine learning · 0.3bayesian network · 0.3kernel methods · 0.2input output kernel regression · 0.2
YearPublicationVenuePosition
2024 MassSpecGym: A benchmark for the discovery and identification of molecules
abstract
The discovery and identification of molecules in biological and environmental samples is crucial for advancing biomedical and chemical sciences. Tandem mass spectrometry (MS/MS) is the leading technique for high-throughput elucidation of molecular structures. However, decoding a molecular structure from its mass spectrum is exceptionally challenging, even when performed by human experts. As a result, the vast majority of acquired MS/MS spectra remain uninterpreted, thereby limiting our understanding of the underlying (bio)chemical processes. Despite decades of progress in machine learning applications for predicting molecular structures from MS/MS spectra, the development of new methods is severely hindered by the lack of standard datasets and evaluation protocols. To address this problem, we propose MassSpecGym -- the first comprehensive benchmark for the discovery and identification of molecules from MS/MS data. Our benchmark comprises the largest publicly available collection of high-quality MS/MS spectra and defines three MS/MS annotation challenges: \textit{de novo} molecular structure generation, molecule retrieval, and spectrum simulation. It includes new evaluation metrics and a generalization-demanding data split, therefore standardizing the MS/MS annotation tasks and rendering the problem accessible to the broad machine learning community. MassSpecGym is publicly available at \url{https://github.com/pluskal-lab/MassSpecGym}.
Roman Bushuiev, Anton Bushuiev, Niek F. de Jonge, Adamo Young, Fleming Kretschmer, Raman Samusevich, Janne Heirman, Fei Wang 0062, Luke Zhang, Kai Dührkop, Marcus Ludwig, Nils A. Haupt, Apurva Kalia, Corinna Brungs, Robin Schmid, Russell Greiner, Bo Wang 0044, David S. Wishart, Liping Liu 0001, Juho Rousu, Wout Bittremieux, Hannes L. Röst, Tytus D. Mak, Soha Hassoun, Florian Huber 0001, Justin J. J. van der Hooft, Michael A. Stravs, Sebastian Böcker, Josef Sivic, Tomás Pluskal
NeurIPS10
2022 Deep kernel learning improves molecular fingerprint prediction from tandem mass spectra
abstract
MOTIVATION: Untargeted metabolomics experiments rely on spectral libraries for structure annotation, but these libraries are vastly incomplete; in silico methods search in structure databases, allowing us to overcome this limitation. The best-performing in silico methods use machine learning to predict a molecular fingerprint from tandem mass spectra, then use the predicted fingerprint to search in a molecular structure database. Predicted molecular fingerprints are also of great interest for compound class annotation, de novo structure elucidation, and other tasks. So far, kernel support vector machines are the best tool for fingerprint prediction. However, they cannot be trained on all publicly available reference spectra because their training time scales cubically with the number of training data. RESULTS: We use the Nyström approximation to transform the kernel into a linear feature map. We evaluate two methods that use this feature map as input: a linear support vector machine and a deep neural network (DNN). For evaluation, we use a cross-validated dataset of 156 017 compounds and three independent datasets with 1734 compounds. We show that the combination of kernel method and DNN outperforms the kernel support vector machine, which is the current gold standard, as well as a DNN on tandem mass spectra on all evaluation datasets. AVAILABILITY AND IMPLEMENTATION: The deep kernel learning method for fingerprint prediction is part of the SIRIUS software, available at https://bio.informatik.uni-jena.de/software/sirius.
Kai Dührkop
Bioinform.1
2018 Heuristic Algorithms for the Maximum Colorful Subtree Problem
abstract
In metabolomics, small molecules are structurally elucidated using tandem mass spectrometry (MS/MS); this computational task can be formulated as the Maximum Colorful Subtree problem, which is NP-hard. Unfortunately, data from a single metabolite requires us to solve hundreds or thousands of instances of this problem - and in a single Liquid Chromatography MS/MS run, hundreds or thousands of metabolites are measured. Here, we comprehensively evaluate the performance of several heuristic algorithms for the problem. Unfortunately, as is often the case in bioinformatics, the structure of the (chemically) true solution is not known to us; therefore we can only evaluate against the optimal solution of an instance. Evaluating the quality of a heuristic based on scores can be misleading: Even a slightly suboptimal solution can be structurally very different from the optimal solution, but it is the structure of a solution and not its score that is relevant for the downstream analysis. To this end, we propose a different evaluation setup: Given a set of candidate instances of which exactly one is known to be correct, the heuristic in question solves each instance to the best of its ability, producing a score for each instance, which is then used to rank the instances. We then evaluate whether the correct instance is ranked highly by the heuristic. We find that one particular heuristic consistently ranks the correct instance in a top position. We also find that the scores of the best heuristic solutions are very close to the optimal score; in contrast, the structure of the solutions can deviate significantly from the optimal structures. Integrating the heuristic allowed us to speed up computations in practice by a factor of 100-fold.
Kai Dührkop, Marie Anne Lataretu, W. Timothy J. White, Sebastian Böcker
WABI1
2018 Bayesian networks for mass spectrometric metabolite identification via molecular fingerprints
abstract
Motivation: Metabolites, small molecules that are involved in cellular reactions, provide a direct functional signature of cellular state. Untargeted metabolomics experiments usually rely on tandem mass spectrometry to identify the thousands of compounds in a biological sample. Recently, we presented CSI:FingerID for searching in molecular structure databases using tandem mass spectrometry data. CSI:FingerID predicts a molecular fingerprint that encodes the structure of the query compound, then uses this to search a molecular structure database such as PubChem. Scoring of the predicted query fingerprint and deterministic target fingerprints is carried out assuming independence between the molecular properties constituting the fingerprint. Results: We present a scoring that takes into account dependencies between molecular properties. As before, we predict posterior probabilities of molecular properties using machine learning. Dependencies between molecular properties are modeled as a Bayesian tree network; the tree structure is estimated on the fly from the instance data. For each edge, we also estimate the expected covariance between the two random variables. For fixed marginal probabilities, we then estimate conditional probabilities using the known covariance. Now, the corrected posterior probability of each candidate can be computed, and candidates are ranked by this score. Modeling dependencies improves identification rates of CSI:FingerID by 2.85 percentage points. Availability and implementation: The new scoring Bayesian (fixed tree) is integrated into SIRIUS 4.0 (https://bio.informatik.uni-jena.de/software/sirius/).
Marcus Ludwig, Kai Dührkop, Sebastian Böcker
Bioinform.2
2016 Fast metabolite identification with Input Output Kernel Regression
abstract
MOTIVATION: An important problematic of metabolomics is to identify metabolites using tandem mass spectrometry data. Machine learning methods have been proposed recently to solve this problem by predicting molecular fingerprint vectors and matching these fingerprints against existing molecular structure databases. In this work we propose to address the metabolite identification problem using a structured output prediction approach. This type of approach is not limited to vector output space and can handle structured output space such as the molecule space. RESULTS: We use the Input Output Kernel Regression method to learn the mapping between tandem mass spectra and molecular structures. The principle of this method is to encode the similarities in the input (spectra) space and the similarities in the output (molecule) space using two kernel functions. This method approximates the spectra-molecule mapping in two phases. The first phase corresponds to a regression problem from the input space to the feature space associated to the output kernel. The second phase is a preimage problem, consisting in mapping back the predicted output feature vectors to the molecule space. We show that our approach achieves state-of-the-art accuracy in metabolite identification. Moreover, our method has the advantage of decreasing the running times for the training step and the test step by several orders of magnitude over the preceding methods. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Céline Brouard, Huibin Shen, Kai Dührkop, Florence d'Alché-Buc, Sebastian Böcker, Juho Rousu
Bioinform.3
2015 Speedy Colorful Subtrees
W. Timothy J. White, Stephan Beyer, Kai Dührkop, Markus Chimani, Sebastian Böcker
COCOON3
2015 Fragmentation Trees Reloaded
Kai Dührkop, Sebastian Böcker
RECOMB1
2014 Metabolite identification through multiple kernel learning on fragmentation trees
abstract
MOTIVATION: Metabolite identification from tandem mass spectrometric data is a key task in metabolomics. Various computational methods have been proposed for the identification of metabolites from tandem mass spectra. Fragmentation tree methods explore the space of possible ways in which the metabolite can fragment, and base the metabolite identification on scoring of these fragmentation trees. Machine learning methods have been used to map mass spectra to molecular fingerprints; predicted fingerprints, in turn, can be used to score candidate molecular structures. RESULTS: Here, we combine fragmentation tree computations with kernel-based machine learning to predict molecular fingerprints and identify molecular structures. We introduce a family of kernels capturing the similarity of fragmentation trees, and combine these kernels using recently proposed multiple kernel learning approaches. Experiments on two large reference datasets show that the new methods significantly improve molecular fingerprint prediction accuracy. These improvements result in better metabolite identification, doubling the number of metabolites ranked at the top position of the candidates list.
Huibin Shen, Kai Dührkop, Sebastian Böcker, Juho Rousu
Bioinform.2
2013 Faster Mass Decomposition
Kai Dührkop, Marcus Ludwig, Marvin Meusel, Sebastian Böcker
WABI1
2012 Fast alignment of fragmentation trees
abstract
MOTIVATION: Mass spectrometry allows sensitive, automated and high-throughput analysis of small molecules such as metabolites. One major bottleneck in metabolomics is the identification of 'unknown' small molecules not in any database. Recently, fragmentation tree alignments have been introduced for the automated comparison of the fragmentation patterns of small molecules. Fragmentation pattern similarities are strongly correlated with the chemical similarity of the molecules, and allow us to cluster compounds based solely on their fragmentation patterns. RESULTS: Aligning fragmentation trees is computationally hard. Nevertheless, we present three exact algorithms for the problem: a dynamic programming (DP) algorithm, a sparse variant of the DP, and an Integer Linear Program (ILP). Evaluation of our methods on three different datasets showed that thousands of alignments can be computed in a matter of minutes using DP, even for 'challenging' instances. Running times of the sparse DP were an order of magnitude better than for the classical DP. The ILP was clearly outperformed by both DP approaches. We also found that for both DP algorithms, computing the 1% slowest alignments required as much time as computing the 99% fastest.
Franziska Hufsky, Kai Dührkop, Florian Rasche, Markus Chimani, Sebastian Böcker
Bioinform.2