VLDB 2026 Research / reviewers in the wild / expert
Jan Krumsiek
dblp:09/2887
· DBLP profile ↗
8ranked-venue papers
3as first author
2since 2021 · last 2022
0000-0003-4734-3791ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
6 papers |
Bioinformatics and computational biology · 63% Medical and health informatics · 19% Computational science and engineering · 19% |
Topics — the 10 heaviest of 12, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology
metabolomics |
1.0 | 2 | 2022 | maplet: an extensible R toolbox for modular and reproducible metabolomics pipelines · Bioinform. 2022 MoDentify: phenotype-driven module identification in metabolomics networks at different resolutions · Bioinform. 2019 |
Medical and health informatics › precision medicine
patient subgroup identification |
0.6 | 1 | 2022 | SGI: automatic clinical subgroup identification in omics datasets · Bioinform. 2022 |
Computational science and engineering › computational reproducibility
reproducible data analysis |
0.6 | 1 | 2022 | maplet: an extensible R toolbox for modular and reproducible metabolomics pipelines · Bioinform. 2022 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis
functional module identification |
0.4 | 1 | 2019 | MoDentify: phenotype-driven module identification in metabolomics networks at different resolutions · Bioinform. 2019 |
Bioinformatics and computational biology
multi-omics data integration |
0.2 | 1 | 2022 | SGI: automatic clinical subgroup identification in omics datasets · Bioinform. 2022 |
Bioinformatics and computational biology › protein analysis › protein-protein interaction
interactome analysis |
0.1 | 1 | 2008 | Bootstrapping the Interactome: Unsupervised Identification of Protein Complexes in Yeast · RECOMB 2008 |
Bioinformatics and computational biology › biological network › network biology
protein complex identification |
0.1 | 1 | 2008 | Bootstrapping the Interactome: Unsupervised Identification of Protein Complexes in Yeast · RECOMB 2008 |
Bioinformatics and computational biology › protein analysis
protein complex prediction |
0.1 | 1 | 2008 | ProCope - protein complex prediction and evaluation · Bioinform. 2008 |
Bioinformatics and computational biology › sequence analysis › sequence visualization
dotplot visualization |
0.1 | 1 | 2007 | Gepard: a rapid and sensitive tool for creating dotplots on genome scale · Bioinform. 2007 |
Bioinformatics and computational biology
sequence analysis |
0.1 | 1 | 2007 | Gepard: a rapid and sensitive tool for creating dotplots on genome scale · Bioinform. 2007 |
Methods — techniques the papers use, named apart from their topics
statistical analysis · 0.6hierarchical clustering · 0.6data visualization · 0.6association testing · 0.6clustering · 0.5correlation network analysis · 0.4unsupervised learning · 0.1interaction scoring · 0.1suffix array · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | SGI: automatic clinical subgroup identification in omics datasetsabstractSUMMARY: The 'Subgroup Identification' (SGI) toolbox provides an algorithm to automatically detect clinical subgroups of samples in large-scale omics datasets. It is based on hierarchical clustering trees in combination with a specifically designed association testing and visualization framework that can process an arbitrary number of clinical parameters and outcomes in a systematic fashion. A multi-block extension allows for the simultaneous use of multiple omics datasets on the same samples. In this article, we first describe the functionality of the toolbox and then demonstrate its capabilities through application examples on a type 2 diabetes metabolomics study as well as two copy number variation datasets from The Cancer Genome Atlas. AVAILABILITY AND IMPLEMENTATION: SGI is an open-source package implemented in R. Package source codes and hands-on tutorials are available at https://github.com/krumsieklab/sgi. The QMdiab metabolomics data is included in the package and can be downloaded from https://doi.org/10.6084/m9.figshare.5904022. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Mustafa Buyukozkan, Karsten Suhre, Jan Krumsiek |
Bioinform. | 3 |
| 2022 | maplet: an extensible R toolbox for modular and reproducible metabolomics pipelinesabstractThis article presents maplet, an open-source R package for the creation of highly customizable, fully reproducible statistical pipelines for metabolomics data analysis. It builds on the SummarizedExperiment data structure to create a centralized pipeline framework for storing data, analysis steps, results and visualizations. maplet's key design feature is its modularity, which offers several advantages, such as ensuring code quality through the maintenance of individual functions and promoting collaborative development by removing technical barriers to code contribution. With over 90 functions, the package includes a wide range of functionalities, covering many widely used statistical approaches and data visualization techniques. AVAILABILITY AND IMPLEMENTATION: The maplet package is implemented in R and freely available at https://github.com/krumsieklab/maplet. Kelsey Chetnik, Elisa Benedetti, Daniel P. Gomari, Annalise Schweickart, Richa Batra, Mustafa Buyukozkan, Matthias Arnold, Jonas Zierer, Karsten Suhre, Jan Krumsiek |
Bioinform. | 11 |
| 2019 | MoDentify: phenotype-driven module identification in metabolomics networks at different resolutionsabstractSummary: Associations of metabolomics data with phenotypic outcomes are expected to span functional modules, which are defined as sets of correlating metabolites that are coordinately regulated. Moreover, these associations occur at different scales, from entire pathways to only a few metabolites; an aspect that has not been addressed by previous methods. Here, we present MoDentify, a free R package to identify regulated modules in metabolomics networks at different layers of resolution. Importantly, MoDentify shows higher statistical power than classical association analysis. Moreover, the package offers direct interactive visualization of the results in Cytoscape. We present an application example using complex, multifluid metabolomics data. Due to its generic character, the method is widely applicable to other types of data. Availability and implementation: https://github.com/krumsieklab/MoDentify (vignette includes detailed workflow). Supplementary information: Supplementary data are available at Bioinformatics online. Kieu Trinh Do, David J. N. P. Rasp, Gabi Kastenmüller, Karsten Suhre, Jan Krumsiek |
Bioinform. | 5 |
| 2012 | On the hypothesis-free testing of metabolite ratios in genome-wide and metabolome-wide association studiesabstractBACKGROUND: Genome-wide association studies (GWAS) with metabolic traits and metabolome-wide association studies (MWAS) with traits of biomedical relevance are powerful tools to identify the contribution of genetic, environmental and lifestyle factors to the etiology of complex diseases. Hypothesis-free testing of ratios between all possible metabolite pairs in GWAS and MWAS has proven to be an innovative approach in the discovery of new biologically meaningful associations. The p-gain statistic was introduced as an ad-hoc measure to determine whether a ratio between two metabolite concentrations carries more information than the two corresponding metabolite concentrations alone. So far, only a rule of thumb was applied to determine the significance of the p-gain. RESULTS: Here we explore the statistical properties of the p-gain through simulation of its density and by sampling of experimental data. We derive critical values of the p-gain for different levels of correlation between metabolite pairs and show that B/(2*α) is a conservative critical value for the p-gain, where α is the level of significance and B the number of tested metabolite pairs. CONCLUSIONS: We show that the p-gain is a well defined measure that can be used to identify statistically significant metabolite ratios in association studies and provide a conservative significance cut-off for the p-gain for use in future association studies with metabolic traits. Ann-Kristin Petersen, Jan Krumsiek, Brigitte Wägele, Fabian J. Theis, Heinz-Erich Wichmann, Christian Gieger, Karsten Suhre |
BMC Bioinform. | 2 |
| 2010 | Odefy -- From discrete to continuous modelsabstractBACKGROUND: Phenomenological information about regulatory interactions is frequently available and can be readily converted to Boolean models. Fully quantitative models, on the other hand, provide detailed insights into the precise dynamics of the underlying system. In order to connect discrete and continuous modeling approaches, methods for the conversion of Boolean systems into systems of ordinary differential equations have been developed recently. As biological interaction networks have steadily grown in size and complexity, a fully automated framework for the conversion process is desirable. RESULTS: We present Odefy, a MATLAB- and Octave-compatible toolbox for the automated transformation of Boolean models into systems of ordinary differential equations. Models can be created from sets of Boolean equations or graph representations of Boolean networks. Alternatively, the user can import Boolean models from the CellNetAnalyzer toolbox, GINSim and the PBN toolbox. The Boolean models are transformed to systems of ordinary differential equations by multivariate polynomial interpolation and optional application of sigmoidal Hill functions. Our toolbox contains basic simulation and visualization functionalities for both, the Boolean as well as the continuous models. For further analyses, models can be exported to SQUAD, GNA, MATLAB script files, the SB toolbox, SBML and R script files. Odefy contains a user-friendly graphical user interface for convenient access to the simulation and exporting functionalities. We illustrate the validity of our transformation approach as well as the usage and benefit of the Odefy toolbox for two biological systems: a mutual inhibitory switch known from stem cell differentiation and a regulatory network giving rise to a specific spatial expression pattern at the mid-hindbrain boundary. CONCLUSIONS: Odefy provides an easy-to-use toolbox for the automatic conversion of Boolean models to systems of ordinary differential equations. It can be efficiently connected to a variety of input and output formats for further analysis and investigations. The toolbox is open-source and can be downloaded at http://cmb.helmholtz-muenchen.de/odefy. Jan Krumsiek, Sebastian Pölsterl, Dominik M. Wittmann, Fabian J. Theis |
BMC Bioinform. | 1 |
| 2008 | Bootstrapping the Interactome: Unsupervised Identification of Protein Complexes in Yeast
Caroline C. Friedel, Jan Krumsiek, Ralf Zimmer |
RECOMB | 2 |
| 2008 | ProCope - protein complex prediction and evaluationabstractSUMMARY: Recent advances in high-throughput technology have increased the quantity of available data on protein complexes and stimulated the development of many new prediction methods. In this article, we present ProCope, a Java software suite for the prediction and evaluation of protein complexes from affinity purification experiments which integrates the major methods for calculating interaction scores and predicting protein complexes published over the last years. Methods can be accessed via a graphical user interface, command line tools and a Java API. Using ProCope, existing algorithms can be applied quickly and reproducibly on new experimental results, individual steps of the different algorithms can be combined in new and innovative ways and new methods can be implemented and integrated in the existing prediction framework. AVAILABILITY: Source code and executables are available at http://www.bio.ifi.lmu.de/Complexes/ProCope/. Jan Krumsiek, Caroline C. Friedel, Ralf Zimmer |
Bioinform. | 1 |
| 2007 | Gepard: a rapid and sensitive tool for creating dotplots on genome scaleabstractUNLABELLED: Gepard provides a user-friendly, interactive application for the quick creation of dotplots. It utilizes suffix arrays to reduce the time complexity of dotplot calculation to Theta(m*log n). A client-server mode, which is a novel feature for dotplot creation software, allows the user to calculate dotplots and color them by functional annotation without any prior downloading of sequence or annotation data. AVAILABILITY: Both source codes and executable binaries are available at http://mips.gsf.de/services/analysis/gepard Jan Krumsiek, Roland Arnold, Thomas Rattei |
Bioinform. | 1 |