VLDB 2026 Research / reviewers in the wild / expert
Armin Rauschenberger
dblp:177/5961
· DBLP profile ↗
6ranked-venue papers
5as first author
5since 2021 · last 2025
0000-0001-6498-4801ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Estimating sparse regression models in multi-task learning and transfer learning through adaptive penalisationabstractMETHOD: Here, we propose a simple two-stage procedure for sharing information between related high-dimensional prediction or classification problems. In both stages, we perform sparse regression separately for each problem. While this is done without prior information in the first stage, we use the coefficients from the first stage as prior information for the second stage. Specifically, we designed feature-specific and sign-specific adaptive weights to share information on feature selection, effect directions, and effect sizes between different problems. RESULTS: The proposed approach is applicable to multi-task learning as well as transfer learning. It provides sparse models (i.e. with few non-zero coefficients for each problem) that are easy to interpret. We show by simulation and application that it tends to select fewer features while achieving a similar predictive performance as compared to available methods. AVAILABILITY AND IMPLEMENTATION: An implementation is available in the R package "sparselink" (https://github.com/rauschenberger/sparselink, https://cran.r-project.org/package=sparselink). Armin Rauschenberger, Petr V. Nazarov, Enrico Glaab |
Bioinform. | 1 |
| 2023 | Penalized regression with multiple sources of prior effectsabstractMOTIVATION: In many high-dimensional prediction or classification tasks, complementary data on the features are available, e.g. prior biological knowledge on (epi)genetic markers. Here we consider tasks with numerical prior information that provide an insight into the importance (weight) and the direction (sign) of the feature effects, e.g. regression coefficients from previous studies. RESULTS: We propose an approach for integrating multiple sources of such prior information into penalized regression. If suitable co-data are available, this improves the predictive performance, as shown by simulation and application. AVAILABILITY AND IMPLEMENTATION: The proposed method is implemented in the R package transreg (https://github.com/lcsb-bds/transreg, https://cran.r-project.org/package=transreg). Armin Rauschenberger, Zied Landoulsi, Mark A. van de Wiel, Enrico Glaab |
Bioinform. | 1 |
| 2022 | Ten quick tips for biomarker discovery and validation analyses using machine learningabstractIntroductionAU : Pleaseconfirmthatallheadinglevelsarerepresentedcorrectly Ramón Díaz-Uriarte, Elisa Gómez de Lope, Rosalba Giugno, Holger Fröhlich, Petr V. Nazarov, Isabel A. Nepomuceno-Chamorro, Armin Rauschenberger, Enrico Glaab |
PLoS Comput. Biol. | 7 |
| 2021 | Predictive and interpretable models via the stacked elastic netabstractMOTIVATION: Machine learning in the biomedical sciences should ideally provide predictive and interpretable models. When predicting outcomes from clinical or molecular features, applied researchers often want to know which features have effects, whether these effects are positive or negative and how strong these effects are. Regression analysis includes this information in the coefficients but typically renders less predictive models than more advanced machine learning techniques. RESULTS: Here, we propose an interpretable meta-learning approach for high-dimensional regression. The elastic net provides a compromise between estimating weak effects for many features and strong effects for some features. It has a mixing parameter to weight between ridge and lasso regularization. Instead of selecting one weighting by tuning, we combine multiple weightings by stacking. We do this in a way that increases predictivity without sacrificing interpretability. AVAILABILITY AND IMPLEMENTATION: The R package starnet is available on GitHub (https://github.com/rauschenberger/starnet) and CRAN (https://CRAN.R-project.org/package=starnet). Armin Rauschenberger, Enrico Glaab, Mark A. van de Wiel |
Bioinform. | 1 |
| 2021 | Predicting correlated outcomes from molecular dataabstractMOTIVATION: Multivariate (multi-target) regression has the potential to outperform univariate (single-target) regression at predicting correlated outcomes, which frequently occur in biomedical and clinical research. Here we implement multivariate lasso and ridge regression using stacked generalization. RESULTS: Our flexible approach leads to predictive and interpretable models in high-dimensional settings, with a single estimate for each input-output effect. In the simulation, we compare the predictive performance of several state-of-the-art methods for multivariate regression. In the application, we use clinical and genomic data to predict multiple motor and non-motor symptoms in Parkinson's disease patients. We conclude that stacked multivariate regression, with our adaptations, is a competitive method for predicting correlated outcomes. AVAILABILITY AND IMPLEMENTATION: The R package joinet is available on GitHub (https://github.com/rauschenberger/joinet) and cran (https://cran.r-project.org/package=joinet). SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Armin Rauschenberger, Enrico Glaab |
Bioinform. | 1 |
| 2016 | Testing for association between RNA-Seq and high-dimensional dataabstractBACKGROUND: Testing for association between RNA-Seq and other genomic data is challenging due to high variability of the former and high dimensionality of the latter. RESULTS: Using the negative binomial distribution and a random-effects model, we develop an omnibus test that overcomes both difficulties. It may be conceptualised as a test of overall significance in regression analysis, where the response variable is overdispersed and the number of explanatory variables exceeds the sample size. CONCLUSIONS: The proposed test can detect genetic and epigenetic alterations that affect gene expression. It can examine complex regulatory mechanisms of gene expression. The R package globalSeq is available from Bioconductor. Armin Rauschenberger, Marianne A. Jonker, Mark A. van de Wiel, Renée X. de Menezes |
BMC Bioinform. | 1 |