VLDB 2026 Research / reviewers in the wild / expert
Richard Röttger
dblp:24/11044
· DBLP profile ↗
11ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0003-4490-5947ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Refinement strategies for Tangram for reliable single-cell to spatial mappingabstractMOTIVATION: Single-cell RNA sequencing (scRNA-seq) provides comprehensive gene expression data at a single-cell level but lacks spatial context. In contrast, spatial transcriptomics captures both spatial and transcriptional information but is limited by resolution, sensitivity, or feasibility. No single technology combines both the high spatial resolution and deep transcriptomic profiling at the single-cell level without tradeoffs. Spatial mapping tools that integrate scRNA-seq and spatial transcriptomics data are crucial to bridge this gap. However, we found that Tangram, one of the most prominent spatial mapping tools, provides inconsistent results over repeated runs. RESULTS: We refine Tangram to achieve more consistent cell mappings and investigate the challenges that arise from data characteristics. We find that the mapping quality depends on the gene expression sparsity. To address this, we (1) train the model on an informative gene subset, (2) apply cell filtering, (3) introduce several forms of regularization, and (4) incorporate neighborhood information. Evaluations on real and simulated mouse datasets demonstrate that this approach improves both gene expression prediction and cell mapping. Consistent cell mapping strengthens the reliability of the projection of cell annotations and features into space, gene imputation, and correction of low-quality measurements. Our pipeline, which includes gene set and hyperparameter selection, can serve as guidance for applying Tangram on other datasets, while our benchmarking framework with data simulation and inconsistency metrics is useful for evaluating other tools or Tangram modifications. AVAILABILITY AND IMPLEMENTATION: The refinements for Tangram and our benchmarking pipeline are available at https://github.com/daisybio/Tangram_Refinement_Strategies. Merle Stahl, Lena J. Straßer, Chit Tong Lio, Judith Bernett, Richard Röttger, Markus List |
Bioinform. | 5 |
| 2024 | Federated singular value decomposition for high-dimensional dataabstractAbstract Federated learning (FL) is emerging as a privacy-aware alternative to classical cloud-based machine learning. In FL, the sensitive data remains in data silos and only aggregated parameters are exchanged. Hospitals and research institutions which are not willing to share their data can join a federated study without breaching confidentiality. In addition to the extreme sensitivity of biomedical data, the high dimensionality poses a challenge in the context of federated genome-wide association studies (GWAS). In this article, we present a federated singular value decomposition algorithm, suitable for the privacy-related and computational requirements of GWAS. Notably, the algorithm has a transmission cost independent of the number of samples and is only weakly dependent on the number of features, because the singular vectors corresponding to the samples are never exchanged and the vectors associated with the features are only transmitted to an aggregator for a fixed number of iterations. Although motivated by GWAS, the algorithm is generically applicable for both horizontally and vertically partitioned data. Anne Hartebrodt, Richard Röttger, David B. Blumenthal |
Data Min. Knowl. Discov. | 2 |
| 2023 | A Comparison of Federated Aggregation Strategies and Architectures for Next-word PredictionabstractFederated learning is an important technique for training language models, which are frequently used for next-word prediction since federated learning allows utilising large quantities of real-life data without compromising the privacy of the data owners. Training a model that generalises well in this setting is a challenging task due to the inherent statistical heterogeneity of the training data, and due to the hardware limitations of private mobile devices. There are different approaches that address these issues, e.g. through model selection, different aggregation and learning strategies, and update compression. In this paper, two popular model architectures, namely Long Short-Term Memory (LSTM) and Gated Recurrent Unit (GRU), are evaluated in centralised and federated settings. For federated learning, the vanilla Federated Averaging algorithm and two alternatives that try to address statistical heterogeneity, namely FedProx, which uses a proximal term to restrict the divergence from the global model during local model training, and Federated Attention, which has similar aims of reducing the distance between models as well to ensure faster convergence and improve generalisation, but is performing this during the aggregation station, are evaluated for their achieved perplexity and accuracy in various settings on two datasets. Based on these results, we provide guidelines on which methods to use, depending on the scenario. Yana Sakhnovych, Richard Röttger, Rudolf Mayer |
IEEE Big Data | 2 |
| 2023 | Human-in-the-Loop Integration with Domain-Knowledge Graphs for Explainable Federated Deep LearningabstractAbstract We explore the integration of domain knowledge graphs into Deep Learning for improved interpretability and explainability using Graph Neural Networks (GNNs). Specifically, a protein-protein interaction (PPI) network is masked over a deep neural network for classification, with patient-specific multi-modal genomic features enriched into the PPI graph’s nodes. Subnetworks that are relevant to the classification (referred to as “disease subnetworks”) are detected using explainable AI. Federated learning is enabled by dividing the knowledge graph into relevant subnetworks, constructing an ensemble classifier, and allowing domain experts to analyze and manipulate detected subnetworks using a developed user interface. Furthermore, the human-in-the-loop principle can be applied with the incorporation of experts, interacting through a sophisticated User Interface (UI) driven by Explainable Artificial Intelligence (xAI) methods, changing the datasets to create counterfactual explanations. The adapted datasets could influence the local model’s characteristics and thereby create a federated version that distils their diverse knowledge in a centralized scenario. This work demonstrates the feasibility of the presented strategies, which were originally envisaged in 2021 and most of it has now been materialized into actionable items. In this paper, we report on some lessons learned during this project. Andreas Holzinger, Anna Saranti, Anne-Christin Hauschild, Jacqueline Michelle Metsch, Dominik Heider, Richard Röttger, Heimo Müller, Jan Baumbach, Bastian Pfeifer |
CD-MAKE | 6 |
| 2023 | The Tower of Babel in Explainable Artificial Intelligence (XAI)abstractAbstract As machine learning (ML) has emerged as the predominant technological paradigm for artificial intelligence (AI), complex black box models such as GPT-4 have gained widespread adoption. Concurrently, explainable AI (XAI) has risen in significance as a counterbalancing force. But the rapid expansion of this research domain has led to a proliferation of terminology and an array of diverse definitions, making it increasingly challenging to maintain coherence. This confusion of languages also stems from the plethora of different perspectives on XAI, e.g. ethics, law, standardization and computer science. This situation threatens to create a “tower of Babel” effect, whereby a multitude of languages impedes the establishment of a common (scientific) ground. In response, this paper first maps different vocabularies, used in ethics, law and standardization. It shows that despite a quest for standardized, uniform XAI definitions, there is still a confusion of languages. Drawing lessons from these viewpoints, it subsequently proposes a methodology for identifying a unified lexicon from a scientific standpoint. This could aid the scientific community in presenting a more unified front to better influence ongoing definition efforts in law and standardization, often without enough scientific representation, which will shape the nature of AI and XAI in the future. David Schneeberger, Richard Röttger, Federico Cabitza, Andrea Campagner, Markus Plass, Heimo Müller, Andreas Holzinger |
CD-MAKE | 2 |
| 2023 | Privacy of Federated QR Decomposition Using Additive Secure Multiparty ComputationabstractFederated learning (FL) is a privacy-aware data mining strategy keeping the private data on the owners’ machine and thereby confidential. The clients compute local models and send them to an aggregator which computes a global model. In hybrid FL, the local parameters are additionally masked using secure aggregation, such that only the global aggregated statistics become available in clear text, not the client specific updates. In this context, we investigate the data leakage of three popular algorithms for QR decomposition, Gram-Schmidt orthonormalization, the Householder algorithm and Givens rotation. We show that, even when using additive SMPC, Givens rotation and the Householder matrix leak raw data and are therefore not suited for this computation paradigm. Gram-Schmidt orthonormalization relies on inner vector products and does not leak raw data points. Anne Hartebrodt, Richard Röttger |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2022 | MS2AI: automated repurposing of public peptide LC-MS data for machine learning applicationsabstractMOTIVATION: Liquid-chromatography mass-spectrometry (LC-MS) is the established standard for analyzing the proteome in biological samples by identification and quantification of thousands of proteins. Machine learning (ML) promises to considerably improve the analysis of the resulting data, however, there is yet to be any tool that mediates the path from raw data to modern ML applications. More specifically, ML applications are currently hampered by three major limitations: (i) absence of balanced training data with large sample size; (ii) unclear definition of sufficiently information-rich data representations for e.g. peptide identification; (iii) lack of benchmarking of ML methods on specific LC-MS problems. RESULTS: We created the MS2AI pipeline that automates the process of gathering vast quantities of MS data for large-scale ML applications. The software retrieves raw data from either in-house sources or from the proteomics identifications database, PRIDE. Subsequently, the raw data are stored in a standardized format amenable for ML, encompassing MS1/MS2 spectra and peptide identifications. This tool bridges the gap between MS and AI, and to this effect we also present an ML application in the form of a convolutional neural network for the identification of oxidized peptides. AVAILABILITY AND IMPLEMENTATION: An open-source implementation of the software can be found at https://gitlab.com/roettgerlab/ms2ai. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Tobias Greisager Rehfeldt, Konrad Krawczyk, Mathias Bøgebjerg, Veit Schwämmle, Richard Röttger |
Bioinform. | 5 |
| 2021 | Federated Principal Component Analysis for Genome-Wide Association StudiesabstractFederated learning (FL) has emerged as a privacy-aware alternative to centralized data analysis, especially for biomedical analyses such as genome-wide association studies (GWAS). The data remains with the owner, which enables studies previously impossible due to privacy protection regulations. Principal component analysis (PCA) is a frequent preprocessing step in GWAS, where the eigenvectors of the sample-by-sample covariance matrix are used as covariates in the statistical tests. Therefore, a federated version of PCA suitable for vertical data partitioning is required for federated GWAS. Existing federated PCA algorithms exchange the complete sample eigenvectors, a potential privacy breach. In this paper, we present a federated PCA algorithm for vertically partitioned data which does not exchange the sample eigenvectors and is hence suitable for federated GWAS. We show that it outperforms existing federated solutions in terms of convergence behavior and scalability. Additionally, we provide a user-friendly privacy-aware web tool to promote acceptance of federated PCA among GWAS researchers. Anne Hartebrodt, Reza Nasirigerdeh, David B. Blumenthal, Richard Röttger |
ICDM | 4 |
| 2017 | Efficient detection of differentially methylated regions using DiMmeRabstractMotivation: Epigenome-wide association studies (EWAS) generate big epidemiological datasets. They aim for detecting differentially methylated DNA regions that are likely to influence transcriptional gene activity and, thus, the regulation of metabolic processes. The by far most widely used technology is the Illumina Methylation BeadChip, which measures the methylation levels of 450 (850) thousand cytosines, in the CpG dinucleotide context in a set of patients compared to a control group. Many bioinformatics tools exist for raw data analysis. However, most of them require some knowledge in the programming language R, have no user interface, and do not offer all necessary steps to guide users from raw data all the way down to statistically significant differentially methylated regions (DMRs) and the associated genes. Results: Here, we present DiMmeR (Discovery of Multiple Differentially Methylated Regions), the first free standalone software that interactively guides with a user-friendly graphical user interface (GUI) scientists the whole way through EWAS data analysis. It offers parallelized statistical methods for efficiently identifying DMRs in both Illumina 450K and 850K EPIC chip data. DiMmeR computes empirical P -values through randomization tests, even for big datasets of hundreds of patients and thousands of permutations within a few minutes on a standard desktop PC. It is independent of any third-party libraries, computes regression coefficients, P -values and empirical P -values, and it corrects for multiple testing. Availability and Implementation: DiMmeR is publicly available at http://dimmer.compbio.sdu.dk . Contact: [email protected]. Supplementary information: Supplementary data are available at Bioinformatics online. Diogo Almeida, Ida Skov, Artur Silva, Fabio Vandin, Qihua Tan, Richard Röttger, Jan Baumbach |
Bioinform. | 6 |
| 2013 | Density parameter estimation for finding clusters of homologous proteins - tracing actinobacterial pathogenicity lifestylesabstractMOTIVATION: Homology detection is a long-standing challenge in computational biology. To tackle this problem, typically all-versus-all BLAST results are coupled with data partitioning approaches resulting in clusters of putative homologous proteins. One of the main problems, however, has been widely neglected: all clustering tools need a density parameter that adjusts the number and size of the clusters. This parameter is crucial but hard to estimate without gold standard data at hand. Developing a gold standard, however, is a difficult and time consuming task. Having a reliable method for detecting clusters of homologous proteins between a huge set of species would open opportunities for better understanding the genetic repertoire of bacteria with different lifestyles. RESULTS: Our main contribution is a method for identifying a suitable and robust density parameter for protein homology detection without a given gold standard. Therefore, we study the core genome of 89 actinobacteria. This allows us to incorporate background knowledge, i.e. the assumption that a set of evolutionarily closely related species should share a comparably high number of evolutionarily conserved proteins (emerging from phylum-specific housekeeping genes). We apply our strategy to find genes/proteins that are specific for certain actinobacterial lifestyles, i.e. different types of pathogenicity. The whole study was performed with transitivity clustering, as it only requires a single intuitive density parameter and has been shown to be well applicable for the task of protein sequence clustering. Note, however, that the presented strategy generally does not depend on our clustering method but can easily be adapted to other clustering approaches. AVAILABILITY: All results are publicly available at http://transclust.mmci.uni-saarland.de/actino_core/ or as Supplementary Material of this article. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Richard Röttger, Prabhav Kalaghatgi, Peng Sun 0008, Siomar de Castro Soares, Vasco Ariston de Carvalho Azevedo, Tobias Wittkop, Jan Baumbach |
Bioinform. | 1 |
| 2012 | How Little Do We Actually Know? On the Size of Gene Regulatory NetworksabstractThe National Center for Biotechnology Information (NCBI) recently announced the availability of whole genome sequences for more than 1,000 species. And the number of sequenced individual organisms is growing. Ongoing improvement of DNA sequencing technology will further contribute to this, enabling large-scale evolution and population genetics studies. However, the availability of sequence information is only the first step in understanding how cells survive, reproduce, and adjust their behavior. The genetic control behind organized development and adaptation of complex organisms still remains widely undetermined. One major molecular control mechanism is transcriptional gene regulation. The direct juxtaposition of the total number of sequenced species to the handful of model organisms with known regulations is surprising. Here, we investigate how little we even know about these model organisms. We aim to predict the sizes of the whole-organism regulatory networks of seven species. In particular, we provide statistical lower bounds for the expected number of regulations. For Escherichia coli we estimate at most 37 percent of the expected gene regulatory interactions to be already discovered, 24 percent for Bacillus subtilis, and <3% human, respectively. We conclude that even for our best researched model organisms we still lack substantial understanding of fundamental molecular control mechanisms, at least on a large scale. Richard Röttger, Ulrich Rückert 0002, Jan Taubert, Jan Baumbach |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |