VLDB 2026 Research / reviewers in the wild / expert
Harald Binder
dblp:05/33
· DBLP profile ↗
23ranked-venue papers
4as first author
5since 2021 · last 2026
0000-0002-5666-8662ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 22 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | mmContext: an open framework for multimodal contrastive learning of omics and text dataabstractSUMMARY: Multimodal approaches are increasingly leveraged for integrating omics data with textual biological knowledge. Yet there is still no accessible, standardized framework that enables systematic comparison of omics representations with different text encoders within a unified workflow. We present mmContext, a lightweight and extensible multimodal embedding framework built on top of the open-source Sentence Transformers library. The software allows researchers to train or apply models that jointly embed omics and text data using any numeric representation stored in an AnnData.obsm layer and any text encoder available in Hugging Face. mmContext supports integration of diverse biological text sources and provides pipelines for training, evaluation, and data preparation. We train and evaluate models for a RNA-Seq and text integration task, and demonstrate their utility through zero-shot classification of cell types and diseases across four independent datasets. By releasing all models, datasets, and tutorials openly, mmContext enables reproducible and accessible multimodal learning for omics-text integration. AVAILABILITY AND IMPLEMENTATION: Pretrained checkpoints and full source code for our custom MMContextEncoder are available on Hugging Face huggingface.co/jo-mengr. The Python package github.com/mengerj/mmcontext provides the model implementation and training and evaluation scripts for custom training. The releases for the publication can be accessed via zenodo: adata_hf_datasets: doi.org/10.5281/zenodo.19185217 and mmContext: doi.org/10.5281/zenodo.19185493. Jonatan Menger, Sonia Maria Krissmer, Clemens Kreutz, Harald Binder, Maren Hackenberg |
Bioinform. | 4 |
| 2025 | Evaluating discrepancies in dimensionality reduction for time-series single-cell RNA-sequencing dataabstractThere are various dimensionality reduction techniques for visually inspecting dynamical patterns in time-series single-cell RNA-sequencing (scRNA-seq) data. However, the lack of one-to-one correspondence between cells across time points makes it difficult to uniquely uncover temporal structure in a low-dimensional manifold. The use of different techniques may thus lead to discrepancies in the representation of dynamical patterns. However, The extent of these discrepancies remains unclear. To investigate this, we propose an approach for reasoning about such discrepancies based on synthetic time-series scRNA-seq data generated by variational autoencoders. The synthetic dynamical patterns induced in a low-dimensional manifold reflect biologically plausible temporal patterns, such as dividing cell clusters during a differentiation process. We consider manifolds from different dimensionality reduction techniques, such as principal component analysis, t-distributed stochastic neighbor embedding, uniform manifold approximation, and projection and single-cell variational inference. We illustrate how the proposed approach allows for reasoning about to what extent low-dimensional manifolds, obtained from different techniques, can capture different dynamical patterns. None of these techniques was found to be consistently superior and the results indicate that they may not reliably represent dynamics when used in isolation, underscoring the need to compare multiple perspectives. Thus, the proposed synthetic dynamical pattern approach provides a foundation for guiding future methods development to detect complex patterns in time-series scRNA-seq data. Maren Hackenberg, Laia Canal Guitart, Rolf Backofen, Harald Binder |
Briefings Bioinform. | 4 |
| 2022 | Stratified neural networks in a time-to-event settingabstractDeep neural networks are frequently employed to predict survival conditional on omics-type biomarkers, e.g., by employing the partial likelihood of Cox proportional hazards model as loss function. Due to the generally limited number of observations in clinical studies, combining different data sets has been proposed to improve learning of network parameters. However, if baseline hazards differ between the studies, the assumptions of Cox proportional hazards model are violated. Based on high dimensional transcriptome profiles from different tumor entities, we demonstrate how using a stratified partial likelihood as loss function allows for accounting for the different baseline hazards in a deep learning framework. Additionally, we compare the partial likelihood with the ranking loss, which is frequently employed as loss function in machine learning approaches due to its seemingly simplicity. Using RNA-seq data from the Cancer Genome Atlas (TCGA) we show that use of stratified loss functions leads to an overall better discriminatory power and lower prediction error compared to their non-stratified counterparts. We investigate which genes are identified to have the greatest marginal impact on prediction of survival when using different loss functions. We find that while similar genes are identified, in particular known prognostic genes receive higher importance from stratified loss functions. Taken together, pooling data from different sources for improved parameter learning of deep neural networks benefits largely from employing stratified loss functions that consider potentially varying baseline hazards. For easy application, we provide PyTorch code for stratified loss functions and an explanatory Jupyter notebook in a GitHub repository. Fabrizio Kuruc, Harald Binder, Moritz Hess |
Briefings Bioinform. | 2 |
| 2021 | Synthetic observations from deep generative models and binary omics data with limited sample sizeabstractDeep generative models can be trained to represent the joint distribution of data, such as measurements of single nucleotide polymorphisms (SNPs) from several individuals. Subsequently, synthetic observations are obtained by drawing from this distribution. This has been shown to be useful for several tasks, such as removal of noise, imputation, for better understanding underlying patterns, or even exchanging data under privacy constraints. Yet, it is still unclear how well these approaches work with limited sample size. We investigate such settings specifically for binary data, e.g. as relevant when considering SNP measurements, and evaluate three frequently employed generative modeling approaches, variational autoencoders (VAEs), deep Boltzmann machines (DBMs) and generative adversarial networks (GANs). This includes conditional approaches, such as when considering gene expression conditional on SNPs. Recovery of pair-wise odds ratios (ORs) is considered as a primary performance criterion. For simulated as well as real SNP data, we observe that DBMs generally can recover structure for up to 300 variables, with a tendency of over-estimating ORs when not carefully tuned. VAEs generally get the direction and relative strength of pairwise relations right, yet with considerable under-estimation of ORs. GANs provide stable results only with larger sample sizes and strong pair-wise relations in the data. Taken together, DBMs and VAEs (in contrast to GANs) appear to be well suited for binary omics data, even at rather small sample sizes. This opens the way for many potential applications where synthetic observations from omics data might be useful. Jens Nußberger, Frederic Boesel, Stefan Lenz, Harald Binder, Moritz Hess |
Briefings Bioinform. | 4 |
| 2021 | Netboost: Boosting-Supported Network Analysis Improves High-Dimensional Omics Prediction in Acute Myeloid Leukemia and Huntington's DiseaseabstractState-of-the art selection methods fail to identify weak but cumulative effects of features found in many high-dimensional omics datasets. Nevertheless, these features play an important role in certain diseases. We present Netboost, a three-step dimension reduction technique. First, a boosting-based filter is combined with the topological overlap measure to identify the essential edges of the network. Second, sparse hierarchical clustering is applied on the selected edges to identify modules and finally module information is aggregated by the first principal components. We demonstrate the application of the newly developed Netboost in combination with CoxBoost for survival prediction of DNA methylation and gene expression data from 180 acute myeloid leukemia (AML) patients and show, based on cross-validated prediction error curve estimates, its prediction superiority over variable selection on the full dataset as well as over an alternative clustering approach. The identified signature related to chromatin modifying enzymes was replicated in an independent dataset, the phase II AMLSG 12-09 study. In a second application we combine Netboost with Random Forest classification and improve the disease classification error in RNA-sequencing data of Huntington's disease mice. Netboost is a freely available Bioconductor R package for dimension reduction and hypothesis generation in high-dimensional omics applications. Pascal Schlosser, Jochen Knaus, Maximilian Schmutz, Konstanze Döhner, Christoph Plass, Lars Bullinger, Rainer Claus, Harald Binder, Michael Lübbert, Martin Schumacher |
IEEE ACM Trans. Comput. Biol. Bioinform. | 8 |
| 2020 | Exploring generative deep learning for omics data using log-linear modelsabstractMOTIVATION: Following many successful applications to image data, deep learning is now also increasingly considered for omics data. In particular, generative deep learning not only provides competitive prediction performance, but also allows for uncovering structure by generating synthetic samples. However, exploration and visualization is not as straightforward as with image applications. RESULTS: We demonstrate how log-linear models, fitted to the generated, synthetic data can be used to extract patterns from omics data, learned by deep generative techniques. Specifically, interactions between latent representations learned by the approaches and generated synthetic data are used to determine sets of joint patterns. Distances of patterns with respect to the distribution of latent representations are then visualized in low-dimensional coordinate systems, e.g. for monitoring training progress. This is illustrated with simulated data and subsequently with cortical single-cell gene expression data. Using different kinds of deep generative techniques, specifically variational autoencoders and deep Boltzmann machines, the proposed approach highlights how the techniques uncover underlying structure. It facilitates the real-world use of such generative deep learning techniques to gain biological insights from omics data. AVAILABILITY AND IMPLEMENTATION: The code for the approach as well as an accompanying Jupyter notebook, which illustrates the application of our approach, is available via the GitHub repository: https://github.com/ssehztirom/Exploring-generative-deep-learning-for-omics-data-by-using-log-linear-models. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Moritz Hess, Maren Hackenberg, Harald Binder |
Bioinform. | 3 |
| 2020 | ideal: an R/Bioconductor package for interactive differential expression analysisabstractBACKGROUND: RNA sequencing (RNA-seq) is an ever increasingly popular tool for transcriptome profiling. A key point to make the best use of the available data is to provide software tools that are easy to use but still provide flexibility and transparency in the adopted methods. Despite the availability of many packages focused on detecting differential expression, a method to streamline this type of bioinformatics analysis in a comprehensive, accessible, and reproducible way is lacking. RESULTS: We developed the ideal software package, which serves as a web application for interactive and reproducible RNA-seq analysis, while producing a wealth of visualizations to facilitate data interpretation. ideal is implemented in R using the Shiny framework, and is fully integrated with the existing core structures of the Bioconductor project. Users can perform the essential steps of the differential expression analysis workflow in an assisted way, and generate a broad spectrum of publication-ready outputs, including diagnostic and summary visualizations in each module, all the way down to functional analysis. ideal also offers the possibility to seamlessly generate a full HTML report for storing and sharing results together with code for reproducibility. CONCLUSION: ideal is distributed as an R package in the Bioconductor project ( http://bioconductor.org/packages/ideal/ ), and provides a solution for performing interactive and reproducible analyses of summarized RNA-seq expression data, empowering researchers with many different profiles (life scientists, clinicians, but also experienced bioinformaticians) to make the ideal use of the data at hand. Federico Marini 0002, Jan Linke, Harald Binder |
BMC Bioinform. | 3 |
| 2019 | pcaExplorer: an R/Bioconductor package for interacting with RNA-seq principal componentsabstractBACKGROUND: Principal component analysis (PCA) is frequently used in genomics applications for quality assessment and exploratory analysis in high-dimensional data, such as RNA sequencing (RNA-seq) gene expression assays. Despite the availability of many software packages developed for this purpose, an interactive and comprehensive interface for performing these operations is lacking. RESULTS: We developed the pcaExplorer software package to enhance commonly performed analysis steps with an interactive and user-friendly application, which provides state saving as well as the automated creation of reproducible reports. pcaExplorer is implemented in R using the Shiny framework and exploits data structures from the open-source Bioconductor project. Users can easily generate a wide variety of publication-ready graphs, while assessing the expression data in the different modules available, including a general overview, dimension reduction on samples and genes, as well as functional interpretation of the principal components. CONCLUSION: pcaExplorer is distributed as an R package in the Bioconductor project ( http://bioconductor.org/packages/pcaExplorer/ ), and is designed to assist a broad range of researchers in the critical step of interactive data exploration. Federico Marini 0002, Harald Binder |
BMC Bioinform. | 2 |
| 2018 | Feasibility of sample size calculation for RNA-seq studiesabstractSample size calculation is a crucial step in study design but is not yet fully established for RNA sequencing (RNA-seq) analyses. To evaluate feasibility and provide guidance, we evaluated RNA-seq sample size tools identified from a systematic search. The focus was on whether real pilot data would be needed for reliable results and on identifying tools that would perform well in scenarios with different levels of biological heterogeneity and fold changes (FCs) between conditions. We used simulations based on real data for tool evaluation. In all settings, the six evaluated tools provided widely different answers, which were strongly affected by FC. Although all tools failed for small FCs, some tools can at least be recommended when closely matching pilot data are available and relatively large FCs are anticipated. Alicia Poplawski, Harald Binder |
Briefings Bioinform. | 2 |
| 2017 | Partitioned learning of deep Boltzmann machines for SNP dataabstractMOTIVATION: Learning the joint distributions of measurements, and in particular identification of an appropriate low-dimensional manifold, has been found to be a powerful ingredient of deep leaning approaches. Yet, such approaches have hardly been applied to single nucleotide polymorphism (SNP) data, probably due to the high number of features typically exceeding the number of studied individuals. RESULTS: After a brief overview of how deep Boltzmann machines (DBMs), a deep learning approach, can be adapted to SNP data in principle, we specifically present a way to alleviate the dimensionality problem by partitioned learning. We propose a sparse regression approach to coarsely screen the joint distribution of SNPs, followed by training several DBMs on SNP partitions that were identified by the screening. Aggregate features representing SNP patterns and the corresponding SNPs are extracted from the DBMs by a combination of statistical tests and sparse regression. In simulated case-control data, we show how this can uncover complex SNP patterns and augment results from univariate approaches, while maintaining type 1 error control. Time-to-event endpoints are considered in an application with acute myeloid leukemia patients, where SNP patterns are modeled after a pre-screening based on gene expression data. The proposed approach identified three SNPs that seem to jointly influence survival in a validation dataset. This indicates the added value of jointly investigating SNPs compared to standard univariate analyses and makes partitioned learning of DBMs an interesting complementary approach when analyzing SNP data. AVAILABILITY AND IMPLEMENTATION: A Julia package is provided at 'http://github.com/binderh/BoltzmannMachines.jl'. CONTACT: [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Moritz Hess, Stefan Lenz, Tamara J. Blätte, Lars Bullinger, Harald Binder |
Bioinform. | 5 |
| 2016 | Systematically evaluating interfaces for RNA-seq analysis from a life scientist perspectiveabstractRNA-sequencing (RNA-seq) has become an established way for measuring gene expression in model organisms and humans. While methods development for refining the corresponding data processing and analysis pipeline is ongoing, protocols for typical steps have been proposed and are widely used. Several user interfaces have been developed for making such analysis steps accessible to life scientists without extensive knowledge of command line tools. We performed a systematic search and evaluation of such interfaces to investigate to what extent these can indeed facilitate RNA-seq data analysis. We found a total of 29 open source interfaces, and six of the more widely used interfaces were evaluated in detail. Central criteria for evaluation were ease of configuration, documentation, usability, computational demand and reporting. No interface scored best in all of these criteria, indicating that the final choice will depend on the specific perspective of users and the corresponding weighting of criteria. Considerable technical hurdles had to be overcome in our evaluation. For many users, this will diminish potential benefits compared with command line tools, leaving room for future improvement of interfaces. Alicia Poplawski, Federico Marini 0002, Moritz Hess, Tanja Zeller, Johanna Mazur, Harald Binder |
Briefings Bioinform. | 6 |
| 2016 | Integrating multiple molecular sources into a clinical risk prediction signature by extracting complementary informationabstractBACKGROUND: High-throughput technology allows for genome-wide measurements at different molecular levels for the same patient, e.g. single nucleotide polymorphisms (SNPs) and gene expression. Correspondingly, it might be beneficial to also integrate complementary information from different molecular levels when building multivariable risk prediction models for a clinical endpoint, such as treatment response or survival. Unfortunately, such a high-dimensional modeling task will often be complicated by a limited overlap of molecular measurements at different levels between patients, i.e. measurements from all molecular levels are available only for a smaller proportion of patients. RESULTS: We propose a sequential strategy for building clinical risk prediction models that integrate genome-wide measurements from two molecular levels in a complementary way. To deal with partial overlap, we develop an imputation approach that allows us to use all available data. This approach is investigated in two acute myeloid leukemia applications combining gene expression with either SNP or DNA methylation data. After obtaining a sparse risk prediction signature e.g. from SNP data, an automatically selected set of prognostic SNPs, by componentwise likelihood-based boosting, imputation is performed for the corresponding linear predictor by a linking model that incorporates e.g. gene expression measurements. The imputed linear predictor is then used for adjustment when building a prognostic signature from the gene expression data. For evaluation, we consider stability, as quantified by inclusion frequencies across resampling data sets. Despite an extremely small overlap in the application example with gene expression and SNPs, several genes are seen to be more stably identified when taking the (imputed) linear predictor from the SNP data into account. In the application with gene expression and DNA methylation, prediction performance with respect to survival also indicates that the proposed approach might work well. CONCLUSIONS: We consider imputation of linear predictor values to be a feasible and sensible approach for dealing with partial overlap in complementary integrative analysis of molecular measurements at different levels. More generally, these results indicate that a complementary strategy for integrating different molecular levels can result in more stable risk prediction signatures, potentially providing a more reliable insight into the underlying biology. Stefanie Hieke, Axel Benner, Richard F. Schlenl, Martin Schumacher, Lars Bullinger, Harald Binder |
BMC Bioinform. | 6 |
| 2015 | Acoustic event source localization for surveillance in reverberant environments supported by an event onset detectionabstractThis contribution presents a robust approach to acoustic event source localization for surveillance under reverberant environmental conditions. In particular, we support the classical generalized cross-correlation algorithm with phase transform weighting (GCC-PHAT) and the steered response power (SRP) algorithm by a sound activity detection and an event onset detector. The proposed algorithmic framework including spatial minimum tracking and smoothing for the suppression of artifacts in the spatial likelihood function significantly outperforms a respective reference approach, decreasing both the miss ratio by up to 9% absolute, and the average angular estimation error by up to 4°. Peter Transfeld, Uwe Martens, Harald Binder, Thomas Schypior, Tim Fingscheidt |
ICASSP | 3 |
| 2015 | Translating bioinformatics in oncology: guilt-by-profiling analysis and identification of KIF18B and CDCA3 as novel driver genes in carcinogenesisabstractMOTIVATION: Co-regulated genes are not identified in traditional microarray analyses, but may theoretically be closely functionally linked [guilt-by-association (GBA), guilt-by-profiling]. Thus, bioinformatics procedures for guilt-by-profiling/association analysis have yet to be applied to large-scale cancer biology. We analyzed 2158 full cancer transcriptomes from 163 diverse cancer entities in regard of their similarity of gene expression, using Pearson's correlation coefficient (CC). Subsequently, 428 highly co-regulated genes (|CC| ≥ 0.8) were clustered unsupervised to obtain small co-regulated networks. A major subnetwork containing 61 closely co-regulated genes showed highly significant enrichment of cancer bio-functions. All genes except kinesin family member 18B (KIF18B) and cell division cycle associated 3 (CDCA3) were of confirmed relevance for tumor biology. Therefore, we independently analyzed their differential regulation in multiple tumors and found severe deregulation in liver, breast, lung, ovarian and kidney cancers, thus proving our GBA hypothesis. Overexpression of KIF18B and CDCA3 in hepatoma cells and subsequent microarray analysis revealed significant deregulation of central cell cycle regulatory genes. Consistently, RT-PCR and proliferation assay confirmed the role of both genes in cell cycle progression. Finally, the prognostic significance of the identified KIF18B- and CDCA3-dependent predictors (P = 0.01, P = 0.04) was demonstrated in three independent HCC cohorts and several other tumors. In summary, we proved the efficacy of large-scale guilt-by-profiling/association strategies in oncology. We identified two novel oncogenes and functionally characterized them. The strong prognostic importance of downstream predictors for HCC and many other tumors indicates the clinical relevance of our findings. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Timo Itzel, Peter Scholz, Thorsten Maass, Markus Krupp, Jens U. Marquardt, Susanne Strand, Diana Becker, Frank Staib, Harald Binder, Stephanie Roessler, Xin Wei Wang 0001, Snorri Thorgeirsson, Martina Müller, Peter R. Galle, Andreas Teufel |
Bioinform. | 9 |
| 2015 | A weighting approach for judging the effect of patient strata on high-dimensional risk prediction signaturesabstractBACKGROUND: High-dimensional molecular measurements, e.g. gene expression data, can be linked to clinical time-to-event endpoints by Cox regression models and regularized estimation approaches, such as componentwise boosting, and can incorporate a large number of covariates as well as provide variable selection. If there is heterogeneity due to known patient subgroups, a stratified Cox model allows for separate baseline hazards in each subgroup. Variable selection will still depend on the relative stratum sizes in the data, which might be a convenience sample and not representative for future applications. Such effects need to be systematically investigated and could even help to more reliably identify components of risk prediction signatures. RESULTS: Correspondingly, we propose a weighted regression approach based on componentwise likelihood-based boosting which is implemented in the R package CoxBoost (https://github.com/binderh/CoxBoost). This approach focuses on building a risk prediction signature for a specific stratum by down-weighting the observations from the other strata using a range of weights. Stability of selection for specific covariates as a function of the weights is investigated by resampling inclusion frequencies, and two types of corresponding visualizations are suggested. This is illustrated for two applications with methylation and gene expression measurements from cancer patients. CONCLUSION: The proposed approach is meant to point out components of risk prediction signatures that are specific to the stratum of interest and components that are also important to other strata. Performance is mostly improved by incorporating down-weighted information from the other strata. This suggests more general usefulness for risk prediction signature development in data with heterogeneity due to known subgroups. Veronika Weyer, Harald Binder |
BMC Bioinform. | 2 |
| 2014 | Combining techniques for screening and evaluating interaction terms on high-dimensional time-to-event dataabstractBACKGROUND: Molecular data, e.g. arising from microarray technology, is often used for predicting survival probabilities of patients. For multivariate risk prediction models on such high-dimensional data, there are established techniques that combine parameter estimation and variable selection. One big challenge is to incorporate interactions into such prediction models. In this feasibility study, we present building blocks for evaluating and incorporating interactions terms in high-dimensional time-to-event settings, especially for settings in which it is computationally too expensive to check all possible interactions. RESULTS: We use a boosting technique for estimation of effects and the following building blocks for pre-selecting interactions: (1) resampling, (2) random forests and (3) orthogonalization as a data pre-processing step. In a simulation study, the strategy that uses all building blocks is able to detect true main effects and interactions with high sensitivity in different kinds of scenarios. The main challenge are interactions composed of variables that do not represent main effects, but our findings are also promising in this regard. Results on real world data illustrate that effect sizes of interactions frequently may not be large enough to improve prediction performance, even though the interactions are potentially of biological relevance. CONCLUSION: Screening interactions through random forests is feasible and useful, when one is interested in finding relevant two-way interactions. The other building blocks also contribute considerably to an enhanced pre-selection of interactions. We determined the limits of interaction detection in terms of necessary effect sizes. Our study emphasizes the importance of making full use of existing methods in addition to establishing new ones. Murat Sariyar, Isabell Hoffmann, Harald Binder |
BMC Bioinform. | 3 |
| 2011 | Graph based fusion of miRNA and mRNA expression data improves clinical outcome prediction in prostate cancerabstractBACKGROUND: One of the main goals in cancer studies including high-throughput microRNA (miRNA) and mRNA data is to find and assess prognostic signatures capable of predicting clinical outcome. Both mRNA and miRNA expression changes in cancer diseases are described to reflect clinical characteristics like staging and prognosis. Furthermore, miRNA abundance can directly affect target transcripts and translation in tumor cells. Prediction models are trained to identify either mRNA or miRNA signatures for patient stratification. With the increasing number of microarray studies collecting mRNA and miRNA from the same patient cohort there is a need for statistical methods to integrate or fuse both kinds of data into one prediction model in order to find a combined signature that improves the prediction. RESULTS: Here, we propose a new method to fuse miRNA and mRNA data into one prediction model. Since miRNAs are known regulators of mRNAs we used the correlations between them as well as the target prediction information to build a bipartite graph representing the relations between miRNAs and mRNAs. This graph was used to guide the feature selection in order to improve the prediction. The method is illustrated on a prostate cancer data set comprising 98 patient samples with miRNA and mRNA expression data. The biochemical relapse was used as clinical endpoint. It could be shown that the bipartite graph in combination with both data sets could improve prediction performance as well as the stability of the feature selection. CONCLUSIONS: Fusion of mRNA and miRNA expression data into one prediction model improves clinical outcome prediction in terms of prediction error and stable feature selection. The R source code of the proposed method is available in the supplement. Stephan Gade, Christine Porzelius, Maria Fälth, Jan C. Brase, Daniela Wuttig, Ruprecht Kuner, Harald Binder, Holger Sültmann, Tim Beißbarth |
BMC Bioinform. | 7 |
| 2009 | Boosting for high-dimensional time-to-event data with competing risksabstractMOTIVATION: For analyzing high-dimensional time-to-event data with competing risks, tailored modeling techniques are required that consider the event of interest and the competing events at the same time, while also dealing with censoring. For low-dimensional settings, proportional hazards models for the subdistribution hazard have been proposed, but an adaptation for high-dimensional settings is missing. In addition, tools for judging the prediction performance of fitted models have to be provided. RESULTS: We propose a boosting approach for fitting proportional subdistribution hazards models for high-dimensional data, that can e.g. incorporate a large number of microarray features, while also taking clinical covariates into account. Prediction performance is evaluated using bootstrap.632+ estimates of prediction error curves, adapted for the competing risks setting. This is illustrated with bladder cancer microarray data, where simultaneous consideration of both, the event of interest and competing events, allows for judging the additional predictive power gained from incorporating microarray measurements. AVAILABILITY: The proposed boosting approach is implemented in the R package CoxBoost and prediction error estimation in the package peperr, both available from CRAN. Harald Binder, Arthur Allignol, Martin Schumacher, Jan Beyersmann |
Bioinform. | 1 |
| 2009 | Parallelized prediction error estimation for evaluation of high-dimensional modelsabstractUNLABELLED: There is a multitude of new techniques that promise to extract predictive information in bioinformatics applications. It has been recognized that a first step for validation of the resulting model fits should rely on proper use of resampling techniques. However, this advice is frequently not followed, potential reasons being difficulty of correct implementation and computational demand. This is addressed by the R package peperr, which is designed for reliable prediction error estimation through resampling, potentially accelerated by parallel execution on a compute cluster. Its interface allows easy connection to newly developed model fitting routines. Performance evaluation of the latter is furthermore guided by diagnostic plots, which helps to detect specific problems due to high-dimensional data structures. AVAILABILITY: http://cran.r-project.org, http://www.imbi.uni-freiburg.de/parallel. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Christine Porzelius, Harald Binder, Martin Schumacher |
Bioinform. | 2 |
| 2009 | Incorporating pathway information into boosting estimation of high-dimensional risk prediction modelsabstractBACKGROUND: There are several techniques for fitting risk prediction models to high-dimensional data, arising from microarrays. However, the biological knowledge about relations between genes is only rarely taken into account. One recent approach incorporates pathway information, available, e.g., from the KEGG database, by augmenting the penalty term in Lasso estimation for continuous response models. RESULTS: As an alternative, we extend componentwise likelihood-based boosting techniques for incorporating pathway information into a larger number of model classes, such as generalized linear models and the Cox proportional hazards model for time-to-event data. In contrast to Lasso-like approaches, no further assumptions for explicitly specifying the penalty structure are needed, as pathway information is incorporated by adapting the penalties for single microarray features in the course of the boosting steps. This is shown to result in improved prediction performance when the coefficients of connected genes have opposite sign. The properties of the fitted models resulting from this approach are then investigated in two application examples with microarray survival data. CONCLUSION: The proposed approach results not only in improved prediction performance but also in structurally different model fits. Incorporating pathway information in the suggested way is therefore seen to be beneficial in several ways. Harald Binder, Martin Schumacher |
BMC Bioinform. | 1 |
| 2008 | Comment on "network-constrained regularization and variable selection for analysis of genomic data"abstractAbstract Contact: [email protected] Supplementary information: Supplementary data are available at Bioinformatics online. Harald Binder, Martin Schumacher |
Bioinform. | 1 |
| 2008 | Allowing for mandatory covariates in boosting estimation of sparse high-dimensional survival modelsabstractBACKGROUND: When predictive survival models are built from high-dimensional data, there are often additional covariates, such as clinical scores, that by all means have to be included into the final model. While there are several techniques for the fitting of sparse high-dimensional survival models by penalized parameter estimation, none allows for explicit consideration of such mandatory covariates. RESULTS: We introduce a new boosting algorithm for censored time-to-event data that shares the favorable properties of existing approaches, i.e., it results in sparse models with good prediction performance, but uses an offset-based update mechanism. The latter allows for tailored penalization of the covariates under consideration. Specifically, unpenalized mandatory covariates can be introduced. Microarray survival data from patients with diffuse large B-cell lymphoma, in combination with the recent, bootstrap-based prediction error curve technique, is used to illustrate the advantages of the new procedure. CONCLUSION: It is demonstrated that it can be highly beneficial in terms of prediction performance to use an estimation procedure that incorporates mandatory covariates into high-dimensional survival models. The new approach also allows to answer the question whether improved predictions are obtained by including microarray features in addition to classical clinical criteria. Harald Binder, Martin Schumacher |
BMC Bioinform. | 1 |
| 2007 | Assessment of survival prediction models based on microarray dataabstractMOTIVATION: In the process of developing risk prediction models, various steps of model building and model selection are involved. If this process is not adequately controlled, overfitting may result in serious overoptimism leading to potentially erroneous conclusions. METHODS: For right censored time-to-event data, we estimate the prediction error for assessing the performance of a risk prediction model (Gerds and Schumacher, 2006; Graf et al., 1999). Furthermore, resampling methods are used to detect overfitting and resulting overoptimism and to adjust the estimates of prediction error (Gerds and Schumacher, 2007). RESULTS: We show how and to what extent the methodology can be used in situations characterized by a large number of potential predictor variables where overfitting may be expected to be overwhelming. This is illustrated by estimating the prediction error of some recently proposed techniques for fitting a multivariate Cox regression model applied to the data of a prognostic study in patients with diffuse large-B-cell lymphoma (DLBCL). AVAILABILITY: Resampling-based estimation of prediction error curves is implemented in an R package called pec available from the authors. Martin Schumacher, Harald Binder, Thomas Gerds |
Bioinform. | 2 |