EDBT 2026 Demo / reviewers in the wild / expert
Jörg Rahnenführer
dblp:40/3957
· DBLP profile ↗
32ranked-venue papers
2as first author
8since 2021 · last 2025
0000-0002-8947-440XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 29 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AlertGS: determining alerts for gene setsabstractMOTIVATION: A typical goal in gene expression studies is identifying certain gene sets enriched with significant genes. The measurement of many gene expression experiments for several concentrations or time points allows the modeling of the concentration/time-response relationship for each gene, and the subsequent estimation of a gene-wise alert. In this work, an approach is proposed to transfer the concept of alerts from single genes to gene sets, yielding a global significance statement and the respective concentration or time where the first enrichment of the gene set can be observed. The methodology is based on a Kolmogorov-Smirnoff type test statistic for each gene set. RESULTS: Simulations show that a majority of these sets can be identified especially for lower numbers of true gene sets with a signal. The false positive rate can be controlled by subsequent decorrelation approaches. Overall, the true gene set-wise alerts are rarely overestimated and rather tend to be underestimated. AVAILABILITY AND IMPLEMENTATION: The code needed to reproduce the simulations and apply the AlertGS methodology is available at the GitHub repository: https://github.com/FKappenberg/AlertGS. Franziska Kappenberg, Jörg Rahnenführer |
Bioinform. | 2 |
| 2023 | Debunking Disinformation with GADMO: A Topic Modeling Analysis of a Comprehensive Corpus of German-language Fact-Checks
Jonas Rieger, Nico Hornig, Jonathan Flossdorf, Henrik Müller, Stephan Mündges, Carsten Jentsch, Jörg Rahnenführer, Christina Elmer |
LDK | 7 |
| 2023 | Designs for the simultaneous inference of concentration-response curvesabstractBACKGROUND: An important problem in toxicology in the context of gene expression data is the simultaneous inference of a large number of concentration-response relationships. The quality of the inference substantially depends on the choice of design of the experiments, in particular, on the set of different concentrations, at which observations are taken for the different genes under consideration. As this set has to be the same for all genes, the efficient planning of such experiments is very challenging. We address this problem by determining efficient designs for the simultaneous inference of a large number of concentration-response models. For that purpose, we both construct a D-optimality criterion for simultaneous inference and a K-means procedure which clusters the support points of the locally D-optimal designs of the individual models. RESULTS: We show that a planning of experiments that addresses the simultaneous inference of a large number of concentration-response relationships yields a substantially more accurate statistical analysis. In particular, we compare the performance of the constructed designs to the ones of other commonly used designs in terms of D-efficiencies and in terms of the quality of the resulting model fits using a real data example dealing with valproic acid. For the quality comparison we perform an extensive simulation study. CONCLUSIONS: The design maximizing the D-optimality criterion for simultaneous inference improves the inference of the different concentration-response relationships substantially. The design based on the K-means procedure also performs well, whereas a log-equidistant design, which was also included in the analysis, performs poorly in terms of the quality of the simultaneous inference. Based on our findings, the D-optimal design for simultaneous inference should be used for upcoming analyses dealing with high-dimensional gene expression data. Leonie Schürmeyer, Kirsten Schorning, Jörg Rahnenführer |
BMC Bioinform. | 3 |
| 2022 | Benchmark of filter methods for feature selection in high-dimensional gene expression survival dataabstractFeature selection is crucial for the analysis of high-dimensional data, but benchmark studies for data with a survival outcome are rare. We compare 14 filter methods for feature selection based on 11 high-dimensional gene expression survival data sets. The aim is to provide guidance on the choice of filter methods for other researchers and practitioners. We analyze the accuracy of predictive models that employ the features selected by the filter methods. Also, we consider the run time, the number of selected features for fitting models with high predictive accuracy as well as the feature selection stability. We conclude that the simple variance filter outperforms all other considered filter methods. This filter selects the features with the largest variance and does not take into account the survival outcome. Also, we identify the correlation-adjusted regression scores filter as a more elaborate alternative that allows fitting models with similar predictive accuracy. Additionally, we investigate the filter methods based on feature rankings, finding groups of similar filters. Andrea Bommert, Thomas Welchowski, Matthias Schmid, Jörg Rahnenführer |
Briefings Bioinform. | 4 |
| 2022 | Correction to: Combining heterogeneous subgroups with graph-structured variable selection priors for Cox regression
Katrin Madjar, Manuela Zucknick, Katja Ickstadt, Jörg Rahnenführer |
BMC Bioinform. | 4 |
| 2021 | Comparison of observation-based and model-based identification of alert concentrations from concentration-expression dataabstractMOTIVATION: An important goal of concentration-response studies in toxicology is to determine an 'alert' concentration where a critical level of the response variable is exceeded. In a classical observation-based approach, only measured concentrations are considered as potential alert concentrations. Alternatively, a parametric curve is fitted to the data that describes the relationship between concentration and response. For a prespecified effect level, both an absolute estimate of the alert concentration and an estimate of the lowest concentration where the effect level is exceeded significantly are of interest. RESULTS: In a simulation study for gene expression data, we compared the observation-based and the model-based approach for both absolute and significant exceedance of the prespecified effect level. Results show that, compared to the observation-based approach, the model-based approach overestimates the true alert concentration less often and more frequently leads to a valid estimate, especially for genes with large variance. AVAILABILITY AND IMPLEMENTATION: The code used for the simulation studies is available via the GitHub repository: https://github.com/FKappenberg/Paper-IdentificationAlertConcentrations. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Franziska Kappenberg, Marianna Grinberg, Xiaoqi Jiang, Annette Kopp-Schneider, Jan G. Hengstler, Jörg Rahnenführer |
Bioinform. | 6 |
| 2021 | Combining heterogeneous subgroups with graph-structured variable selection priors for Cox regressionabstractBACKGROUND: Important objectives in cancer research are the prediction of a patient's risk based on molecular measurements such as gene expression data and the identification of new prognostic biomarkers (e.g. genes). In clinical practice, this is often challenging because patient cohorts are typically small and can be heterogeneous. In classical subgroup analysis, a separate prediction model is fitted using only the data of one specific cohort. However, this can lead to a loss of power when the sample size is small. Simple pooling of all cohorts, on the other hand, can lead to biased results, especially when the cohorts are heterogeneous. RESULTS: We propose a new Bayesian approach suitable for continuous molecular measurements and survival outcome that identifies the important predictors and provides a separate risk prediction model for each cohort. It allows sharing information between cohorts to increase power by assuming a graph linking predictors within and across different cohorts. The graph helps to identify pathways of functionally related genes and genes that are simultaneously prognostic in different cohorts. CONCLUSIONS: Results demonstrate that our proposed approach is superior to the standard approaches in terms of prediction performance and increased power in variable selection when the sample size is small. Katrin Madjar, Manuela Zucknick, Katja Ickstadt, Jörg Rahnenführer |
BMC Bioinform. | 4 |
| 2021 | MODES: model-based optimization on distributed embedded systemsabstractAbstract The predictive performance of a machine learning model highly depends on the corresponding hyper-parameter setting. Hence, hyper-parameter tuning is often indispensable. Normally such tuning requires the dedicated machine learning model to be trained and evaluated on centralized data to obtain a performance estimate. However, in a distributed machine learning scenario, it is not always possible to collect all the data from all nodes due to privacy concerns or storage limitations. Moreover, if data has to be transferred through low bandwidth connections it reduces the time available for tuning. Model-Based Optimization (MBO) is one state-of-the-art method for tuning hyper-parameters but the application on distributed machine learning models or federated learning lacks research. This work proposes a framework $$\textit{MODES}$$ MODES that allows to deploy MBO on resource-constrained distributed embedded systems. Each node trains an individual model based on its local data. The goal is to optimize the combined prediction accuracy. The presented framework offers two optimization modes: (1) $$\textit{MODES}$$ MODES -B considers the whole ensemble as a single black box and optimizes the hyper-parameters of each individual model jointly, and (2) $$\textit{MODES}$$ MODES -I considers all models as clones of the same black box which allows it to efficiently parallelize the optimization in a distributed setting. We evaluate $$\textit{MODES}$$ MODES by conducting experiments on the optimization for the hyper-parameters of a random forest and a multi-layer perceptron. The experimental results demonstrate that, with an improvement in terms of mean accuracy ( $$\textit{MODES}$$ MODES -B), run-time efficiency ( $$\textit{MODES}$$ MODES -I), and statistical stability for both modes, $$\textit{MODES}$$ MODES outperforms the baseline, i.e., carry out tuning with MBO on each node individually with its local sub-data set. Jiang Bian 0003, Jakob Richter, Kuan-Hsun Chen, Jörg Rahnenführer, Haoyi Xiong, Jian-Jia Chen |
Mach. Learn. | 5 |
| 2020 | Model-based optimization with concept driftsabstractModel-based Optimization (MBO) is a method to optimize expensive black-box functions that uses a surrogate to guide the search. We propose two practical approaches that allow MBO to optimize black-box functions where the relation between input and output changes over time, which are known as dynamic optimization problems (DOPs). The window approach trains the surrogate only on the most recent observations, and the time-as-covariate approach includes the time as an additional input variable in the surrogate, giving it the ability to learn the effect of the time on the outcomes. We focus on problems where the change happens systematically and label this systematic change concept drift. To benchmark our methods we define a set of benchmark functions built from established synthetic static functions that are extended with controlled drifts. We evaluate how the proposed approaches handle scenarios of no drift, sudden drift and incremental drift. The results show that both new methods improve the performance if a drift is present. For higher-dimensional multimodal problems the window approach works best and on lower-dimensional problems, where it is easier for the surrogate to capture the influence of the time, the time-as-covariate approach works better. Jakob Richter, Jian-Jia Chen, Jörg Rahnenführer, Michel Lang |
GECCO | 4 |
| 2020 | Improving Latent Dirichlet Allocation: On Reliability of the Novel Method LDAPrototype
Jonas Rieger, Jörg Rahnenführer, Carsten Jentsch |
NLDB | 2 |
| 2020 | Cost-Constrained feature selection in binary classification: adaptations for greedy forward selection and genetic algorithmsabstractBACKGROUND: With modern methods in biotechnology, the search for biomarkers has advanced to a challenging statistical task exploring high dimensional data sets. Feature selection is a widely researched preprocessing step to handle huge numbers of biomarker candidates and has special importance for the analysis of biomedical data. Such data sets often include many input features not related to the diagnostic or therapeutic target variable. A less researched, but also relevant aspect for medical applications are costs of different biomarker candidates. These costs are often financial costs, but can also refer to other aspects, for example the decision between a painful biopsy marker and a simple urine test. In this paper, we propose extensions to two feature selection methods to control the total amount of such costs: greedy forward selection and genetic algorithms. In comprehensive simulation studies of binary classification tasks, we compare the predictive performance, the run-time and the detection rate of relevant features for the new proposed methods and five baseline alternatives to handle budget constraints. RESULTS: In simulations with a predefined budget constraint, our proposed methods outperform the baseline alternatives, with just minor differences between them. Only in the scenario without an actual budget constraint, our adapted greedy forward selection approach showed a clear drop in performance compared to the other methods. However, introducing a hyperparameter to adapt the benefit-cost trade-off in this method could overcome this weakness. CONCLUSIONS: In feature cost scenarios, where a total budget has to be met, common feature selection algorithms are often not suitable to identify well performing subsets for a modelling task. Adaptations of these algorithms such as the ones proposed in this paper can help to tackle this problem. Rudolf Jagdhuber, Michel Lang, Arnulf Stenzl, Jochen Neuhaus, Jörg Rahnenführer |
BMC Bioinform. | 5 |
| 2019 | Model-based optimization of subgroup weights for survival analysisabstractMOTIVATION: To obtain a reliable prediction model for a specific cancer subgroup or cohort is often difficult due to limited sample size and, in survival analysis, due to potentially high censoring rates. Sometimes similar data from other patient subgroups are available, e.g. from other clinical centers. Simple pooling of all subgroups can decrease the variance of the predicted parameters of the prediction models, but also increase the bias due to heterogeneity between the cohorts. A promising compromise is to identify those subgroups with a similar relationship between covariates and target variable and then include only these for model building. RESULTS: We propose a subgroup-based weighted likelihood approach for survival prediction with high-dimensional genetic covariates. When predicting survival for a specific subgroup, for every other subgroup an individual weight determines the strength with which its observations enter into model building. MBO (model-based optimization) can be used to quickly find a good prediction model in the presence of a large number of hyperparameters. We use MBO to identify the best model for survival prediction of a specific subgroup by optimizing the weights for additional subgroups for a Cox model. The approach is evaluated on a set of lung cancer cohorts with gene expression measurements. The resulting models have competitive prediction quality, and they reflect the similarity of the corresponding cancer subgroups, with both weights close to 0 and close to 1 and medium weights. AVAILABILITY AND IMPLEMENTATION: mlrMBO is implemented as an R-package and is freely available at http://github.com/mlr-org/mlrMBO. Jakob Richter, Katrin Madjar, Jörg Rahnenführer |
Bioinform. | 3 |
| 2017 | Variable selection for disease progression models: methods for oncogenetic trees and application to cancer and HIVabstractBACKGROUND: Disease progression models are important for understanding the critical steps during the development of diseases. The models are imbedded in a statistical framework to deal with random variations due to biology and the sampling process when observing only a finite population. Conditional probabilities are used to describe dependencies between events that characterise the critical steps in the disease process. Many different model classes have been proposed in the literature, from simple path models to complex Bayesian networks. A popular and easy to understand but yet flexible model class are oncogenetic trees. These have been applied to describe the accumulation of genetic aberrations in cancer and HIV data. However, the number of potentially relevant aberrations is often by far larger than the maximal number of events that can be used for reliably estimating the progression models. Still, there are only a few approaches to variable selection, which have not yet been investigated in detail. RESULTS: We fill this gap and propose specifically for oncogenetic trees ten variable selection methods, some of these being completely new. We compare them in an extensive simulation study and on real data from cancer and HIV. It turns out that the preselection of events by clique identification algorithms performs best. Here, events are selected if they belong to the largest or the maximum weight subgraph in which all pairs of vertices are connected. CONCLUSIONS: The variable selection method of identifying cliques finds both the important frequent events and those related to disease pathways. Katrin Hainke, Sebastian Szugat, Roland Fried, Jörg Rahnenführer |
BMC Bioinform. | 4 |
| 2017 | DISMS2: A flexible algorithm for direct proteome- wide distance calculation of LC-MS/MS runsabstractBACKGROUND: The classification of samples on a molecular level has manifold applications, from patient classification regarding cancer treatment to phylogenetics for identifying evolutionary relationships between species. Modern methods employ the alignment of DNA or amino acid sequences, mostly not genome-wide but only on selected parts of the genome. Recently proteomics-based approaches have become popular. An established method for the identification of peptides and proteins is liquid chromatography-tandem mass spectrometry (LC-MS/MS). First, protein sequences from MS/MS spectra are identified by means of database searches, given samples with known genome-wide sequence information, then sequence based methods are applied. Alternatively, de novo peptide sequencing algorithms annotate MS/MS spectra and deduce peptide/protein information without a database. A newer approach independent of additional information is to directly compare unidentified tandem mass spectra. The challenge then is to compute the distance between pairwise MS/MS runs consisting of thousands of spectra. METHODS: We present DISMS2, a new algorithm to calculate proteome-wide distances directly from MS/MS data, extending the algorithm compareMS2, an approach that also uses a spectral comparison pipeline. RESULTS: Our new more flexible algorithm, DISMS2, allows for the choice of the spectrum distance measure and includes different spectra preprocessing and filtering steps that can be tailored to specific situations by parameter optimization. CONCLUSIONS: DISMS2 performs well for samples from species with and without database annotation and thus has clear advantages over methods that are purely based on database search. Vera Rieder, Bernhard Blank-Landeshammer, Marleen Stuhr, Tilman Schell, Karsten Biß, Laxmikanth Kollipara, Achim Meyer, Markus Pfenninger, Hildegard Westphal, Albert Sickmann, Jörg Rahnenführer |
BMC Bioinform. | 11 |
| 2016 | TiMEx: a waiting time model for mutually exclusive cancer alterationsabstractMOTIVATION: Despite recent technological advances in genomic sciences, our understanding of cancer progression and its driving genetic alterations remains incomplete. RESULTS: We introduce TiMEx, a generative probabilistic model for detecting patterns of various degrees of mutual exclusivity across genetic alterations, which can indicate pathways involved in cancer progression. TiMEx explicitly accounts for the temporal interplay between the waiting times to alterations and the observation time. In simulation studies, we show that our model outperforms previous methods for detecting mutual exclusivity. On large-scale biological datasets, TiMEx identifies gene groups with strong functional biological relevance, while also proposing new candidates for biological validation. TiMEx possesses several advantages over previous methods, including a novel generative probabilistic model of tumorigenesis, direct estimation of the probability of mutual exclusivity interaction, computational efficiency and high sensitivity in detecting gene groups involving low-frequency alterations. AVAILABILITY AND IMPLEMENTATION: TiMEx is available as a Bioconductor R package at www.bsse.ethz.ch/cbg/software/TiMEx CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Simona Constantinescu, Ewa Szczurek, Pejman Mohammadi 0001, Jörg Rahnenführer, Niko Beerenwinkel |
Bioinform. | 4 |
| 2011 | Survival models with preclustered gene groups as covariatesabstractBACKGROUND: An important application of high dimensional gene expression measurements is the risk prediction and the interpretation of the variables in the resulting survival models. A major problem in this context is the typically large number of genes compared to the number of observations (individuals). Feature selection procedures can generate predictive models with high prediction accuracy and at the same time low model complexity. However, interpretability of the resulting models is still limited due to little knowledge on many of the remaining selected genes. Thus, we summarize genes as gene groups defined by the hierarchically structured Gene Ontology (GO) and include these gene groups as covariates in the hazard regression models. Since expression profiles within GO groups are often heterogeneous, we present a new method to obtain subgroups with coherent patterns. We apply preclustering to genes within GO groups according to the correlation of their gene expression measurements. RESULTS: We compare Cox models for modeling disease free survival times of breast cancer patients. Besides classical clinical covariates we consider genes, GO groups and preclustered GO groups as additional genomic covariates. Survival models with preclustered gene groups as covariates have similar prediction accuracy as models built only with single genes or GO groups. CONCLUSIONS: The preclustering information enables a more detailed analysis of the biological meaning of covariates selected in the final models. Compared to models built only with single genes there is additional functional information contained in the GO annotation, and compared to models using GO groups as covariates the preclustering yields coherent representative gene expression profiles. Kai Kammers, Michel Lang, Jan G. Hengstler, Marcus Schmidt, Jörg Rahnenführer |
BMC Bioinform. | 5 |
| 2010 | Going from where to why - interpretable prediction of protein subcellular localizationabstractMOTIVATION: Protein subcellular localization is pivotal in understanding a protein's function. Computational prediction of subcellular localization has become a viable alternative to experimental approaches. While current machine learning-based methods yield good prediction accuracy, most of them suffer from two key problems: lack of interpretability and dealing with multiple locations. RESULTS: We present YLoc, a novel method for predicting protein subcellular localization that addresses these issues. Due to its simple architecture, YLoc can identify the relevant features of a protein sequence contributing to its subcellular localization, e.g. localization signals or motifs relevant to protein sorting. We present several example applications where YLoc identifies the sequence features responsible for protein localization, and thus reveals not only to which location a protein is transported to, but also why it is transported there. YLoc also provides a confidence estimate for the prediction. Thus, the user can decide what level of error is acceptable for a prediction. Due to a probabilistic approach and the use of several thousands of dual-targeted proteins, YLoc is able to predict multiple locations per protein. YLoc was benchmarked using several independent datasets for protein subcellular localization and performs on par with other state-of-the-art predictors. Disregarding low-confidence predictions, YLoc can achieve prediction accuracies of over 90%. Moreover, we show that YLoc is able to reliably predict multiple locations and outperforms the best predictors in this area. AVAILABILITY: www.multiloc.org/YLoc. Sebastian Briesemeister, Jörg Rahnenführer, Oliver Kohlbacher |
Bioinform. | 2 |
| 2010 | Comparison of scores for bimodality of gene expression distributions and genome-wide evaluation of the prognostic relevance of high-scoring genesabstractBACKGROUND: A major goal of the analysis of high-dimensional RNA expression data from tumor tissue is to identify prognostic signatures for discriminating patient subgroups. For this purpose genome-wide identification of bimodally expressed genes from gene array data is relevant because distinguishability of high and low expression groups is easier compared to genes with unimodal expression distributions.Recently, several methods for the identification of genes with bimodal distributions have been introduced. A straightforward approach is to cluster the expression values and score the distance between the two distributions. Other scores directly measure properties of the distribution. The kurtosis, e.g., measures divergence from a normal distribution. An alternative is the outlier-sum statistic that identifies genes with extremely high or low expression values in a subset of the samples. RESULTS: We compare and discuss scores for bimodality for expression data. For the genome-wide identification of bimodal genes we apply all scores to expression data from 194 patients with node-negative breast cancer. Further, we present the first comprehensive genome-wide evaluation of the prognostic relevance of bimodal genes. We first rank genes according to bimodality scores and define two patient subgroups based on expression values. Then we assess the prognostic significance of the top ranking bimodal genes by comparing the survival functions of the two patient subgroups. We also evaluate the global association between the bimodal shape of expression distributions and survival times with an enrichment type analysis.Various cluster-based methods lead to a significant overrepresentation of prognostic genes. A striking result is obtained with the outlier-sum statistic (p < 10-12). Many genes with heavy tails generate subgroups of patients with different prognosis. CONCLUSIONS: Genes with high bimodality scores are promising candidates for defining prognostic patient subgroups from expression data. We discuss advantages and disadvantages of the different scores for prognostic purposes. The outlier-sum statistic may be particularly valuable for the identification of genes to be included in prognostic signatures. Among the genes identified as bimodal in the breast cancer data set several have not yet previously been recognized to be prognostic and bimodally expressed in breast cancer. Birte Hellwig, Jan G. Hengstler, Marcus Schmidt, Mathias C. Gehrmann, Wiebke Schormann, Jörg Rahnenführer |
BMC Bioinform. | 6 |
| 2009 | Retention time alignment algorithms for LC/MS data must consider non-linear shiftsabstractMOTIVATION: Proteomics has particularly evolved to become of high interest for the field of biomarker discovery and drug development. Especially the combination of liquid chromatography and mass spectrometry (LC/MS) has proven to be a powerful technique for analyzing protein mixtures. Clinically orientated proteomic studies will have to compare hundreds of LC/MS runs at a time. In order to compare different runs, sophisticated preprocessing steps have to be performed. An important step is the retention time (rt) alignment of LC/MS runs. Especially non-linear shifts in the rt between pairs of LC/MS runs make this a crucial and non-trivial problem. RESULTS: For the purpose of demonstrating the particular importance of correcting non-linear rt shifts, we evaluate and compare different alignment algorithms. We present and analyze two versions of a new algorithm that is based on regression techniques, once assuming and estimating only linear shifts and once also allowing for the estimation of non-linear shifts. As an example for another type of alignment method we use an established alignment algorithm based on shifting vectors that we adapted to allow for correcting non-linear shifts also. In a simulation study, we show that rt alignment procedures that can estimate non-linear shifts yield clearly better alignments. This is even true under mild non-linear deviations. AVAILABILITY: R code for the regression-based alignment methods and simulated datasets are available at http://www.statistik.tu-dortmund.de/genetik-publikationen-alignment.html. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Katharina Podwojski, Arno Fritsch, Daniel C. Chamrad, Wolfgang Paul 0001, Barbara Sitek, Kai Stühler, Petra Mutzel, Christian Stephan, Helmut E. Meyer, Wolfgang Urfer, Katja Ickstadt, Jörg Rahnenführer |
Bioinform. | 12 |
| 2008 | Rtreemix: an R package for estimating evolutionary pathways and genetic progression scoresabstractAbstract Summary: In genetics, many evolutionary pathways can be modeled by the ordered accumulation of permanent changes. Mixture models of mutagenetic trees have been used to describe disease progression in cancer and in HIV. In cancer, progression is modeled by the accumulation of chromosomal gains and losses in tumor cells; in HIV, the accumulation of drug resistance-associated mutations in the viral genome is known to be associated with disease progression. From such evolutionary models, genetic progression scores can be derived that assign measures for the disease state to single patients. Rtreemix is an R package for estimating mixture models of evolutionary pathways from observed cross-sectional data and for estimating associated genetic progression scores. The package also provides extended functionality for estimating confidence intervals for estimated model parameters and for evaluating the stability of the estimated evolutionary mixture models. Availability: Rtreemix is an R package that is freely available from the Bioconductor project at http://www.bioconductor.org and runs on Linux and Windows. Contact: [email protected] Jasmina Bogojeska, Adrian Alexa, André Altmann, Thomas Lengauer, Jörg Rahnenführer |
Bioinform. | 5 |
| 2008 | Stability analysis of mixtures of mutagenetic treesabstractBACKGROUND: Mixture models of mutagenetic trees are evolutionary models that capture several pathways of ordered accumulation of genetic events observed in different subsets of patients. They were used to model HIV progression by accumulation of resistance mutations in the viral genome under drug pressure and cancer progression by accumulation of chromosomal aberrations in tumor cells. From the mixture models a genetic progression score (GPS) can be derived that estimates the genetic status of single patients according to the corresponding progression along the tree models. GPS values were shown to have predictive power for estimating drug resistance in HIV or the survival time in cancer. Still, the reliability of the exact values of such complex markers derived from graphical models can be questioned. RESULTS: In a simulation study, we analyzed various aspects of the stability of estimated mutagenetic trees mixture models. It turned out that the induced probabilistic distributions and the tree topologies are recovered with high precision by an EM-like learning algorithm. However, only for models with just one major model component, also GPS values of single patients can be reliably estimated. CONCLUSION: It is encouraging that the estimation process of mutagenetic trees mixture models can be performed with high confidence regarding induced probability distributions and the general shape of the tree topologies. For a model with only one major disease progression process, even genetic progression scores for single patients can be reliably estimated. However, for models with more than one relevant component, alternative measures should be introduced for estimating the stage of disease progression. Jasmina Bogojeska, Thomas Lengauer, Jörg Rahnenführer |
BMC Bioinform. | 3 |
| 2007 | Conformational analysis of alternative protein structuresabstractMOTIVATION: Alternative structural models determined experimentally are available for an increasing number of proteins. Structural and functional studies of these proteins need to take these models into consideration as they can present considerable structural differences. The characterization of the structural differences and similarities between these models is a fundamental task in structural biology requiring appropriate methods. RESULTS: We propose a method for characterizing sets of alternative structural models. Three types of analysis are performed: grouping according to structural similarity, visualization and detection of structural variation and comparison of subsets for identifying and locating distinct conformational states. The alpha carbon atoms are used in order to analyse the backbone conformations. Alternatively, side-chain atoms are used for detailed conformational analysis of specific sites. The method takes into account estimates of atom coordinate uncertainty. The invariant regions are used to generate optimal superpositions of these models. We present the results obtained for three proteins showing different degrees of conformational variability: relative motion of two structurally conserved subdomains, a disordered subdomain and flexibility in the functional site associated with ligand binding. The method has been applied in the analysis of the alternative models available in SCOP. Considerable structural variability can be observed for most proteins. AVAILABILITY: The results of the analysis of the SCOP alternative models, the estimates of coordinate uncertainty as well as the source code of the implementation are available in the STRuster web site: http://struster.bioinf.mpi-inf.mpg.de. Francisco S. Domingues, Jörg Rahnenführer, Thomas Lengauer |
Bioinform. | 2 |
| 2006 | Improved scoring of functional groups from gene expression data by decorrelating GO graph structureabstractMOTIVATION: The result of a typical microarray experiment is a long list of genes with corresponding expression measurements. This list is only the starting point for a meaningful biological interpretation. Modern methods identify relevant biological processes or functions from gene expression data by scoring the statistical significance of predefined functional gene groups, e.g. based on Gene Ontology (GO). We develop methods that increase the explanatory power of this approach by integrating knowledge about relationships between the GO terms into the calculation of the statistical significance. RESULTS: We present two novel algorithms that improve GO group scoring using the underlying GO graph topology. The algorithms are evaluated on real and simulated gene expression data. We show that both methods eliminate local dependencies between GO terms and point to relevant areas in the GO graph that remain undetected with state-of-the-art algorithms for scoring functional terms. A simulation study demonstrates that the new methods exhibit a higher level of detecting relevant biological terms than competing methods. Adrian Alexa, Jörg Rahnenführer, Thomas Lengauer |
Bioinform. | 2 |
| 2006 | A new measure for functional similarity of gene products based on Gene OntologyabstractBACKGROUND: Gene Ontology (GO) is a standard vocabulary of functional terms and allows for coherent annotation of gene products. These annotations provide a basis for new methods that compare gene products regarding their molecular function and biological role. RESULTS: We present a new method for comparing sets of GO terms and for assessing the functional similarity of gene products. The method relies on two semantic similarity measures; simRel and funSim. One measure (simRel) is applied in the comparison of the biological processes found in different groups of organisms. The other measure (funSim) is used to find functionally related gene products within the same or between different genomes. Results indicate that the method, in addition to being in good agreement with established sequence similarity approaches, also provides a means for the identification of functionally related proteins independent of evolutionary relationships. The method is also applied to estimating functional similarity between all proteins in Saccharomyces cerevisiae and to visualizing the molecular function space of yeast in a map of the functional space. A similar approach is used to visualize the functional relationships between protein families. CONCLUSION: The approach enables the comparison of the underlying molecular biology of different taxonomic groups and provides a new comparative genomics tool identifying functionally related gene products independent of homology. The proposed map of the functional space provides a new global view on the functional relationships between gene products or protein families. Andreas Schlicker, Francisco S. Domingues, Jörg Rahnenführer, Thomas Lengauer |
BMC Bioinform. | 3 |
| 2005 | Mtreemix: a software package for learning and using mixture models of mutagenetic treesabstractSUMMARY: Mixture models of mutagenetic trees constitute a class of probabilistic models for describing evolutionary processes that are characterized by the accumulation of permanent genetic changes. They have been applied to model the accumulation of chromosomal gains and losses in tumor development and the development of drug resistance-associated mutations in the HIV genome.Mtreemix is a software package for estimating mutagenetic trees mixture models from observed cross-sectional data and for using these models for predictions. We provide programs for model fitting, model selection, simulation, likelihood computation and waiting time estimation. AVAILABILITY: Mtreemix, including source code, documentation, sample data files and precompiled Solaris and Linux binaries, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de/ Niko Beerenwinkel, Jörg Rahnenführer, Rolf Kaiser, Daniel Hoffmann, Joachim Selbig, Thomas Lengauer |
Bioinform. | 2 |
| 2005 | Computational methods for the design of effective therapies against drug resistant HIV strainsabstractThe development of drug resistance is a major obstacle to successful treatment of HIV infection. The extraordinary replication dynamics of HIV facilitates its escape from selective pressure exerted by the human immune system and by combination drug therapy. We have developed several computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genomic data. Niko Beerenwinkel, Tobias Sing, Thomas Lengauer, Jörg Rahnenführer, Kirsten Roomp, Igor Savenkov, Roman Fischer, Daniel Hoffmann, Joachim Selbig, Klaus Korn, Hauke Walter, Thomas Berg, Patrick Braun, Gerd Fätkenheuer, Mark Oette, Jürgen K. Rockstroh, Bernd Kupfer, Rolf Kaiser, Martin Däumer |
Bioinform. | 4 |
| 2005 | Estimating cancer survival and clinical outcome based on genetic tumor progression scoresabstractMOTIVATION: In cancer research, prediction of time to death or relapse is important for a meaningful tumor classification and selecting appropriate therapies. Survival prognosis is typically based on clinical and histological parameters. There is increasing interest in identifying genetic markers that better capture the status of a tumor in order to improve on existing predictions. The accumulation of genetic alterations during tumor progression can be used for the assessment of the genetic status of the tumor. For modeling dependences between the genetic events, evolutionary tree models have been applied. RESULTS: Mixture models of oncogenetic trees provide a probabilistic framework for the estimation of typical pathogenetic routes. From these models we derive a genetic progression score (GPS) that estimates the genetic status of a tumor. GPS is calculated for glioblastoma patients from loss of heterozygosity measurements and for prostate cancer patients from comparative genomic hybridization measurements. Cox proportional hazard models are then fitted to observed survival times of glioblastoma patients and to times until PSA relapse following radical prostatectomy of prostate cancer patients. It turns out that the genetically defined GPS is predictive even after adjustment for classical clinical markers and thus can be considered a medically relevant prognostic factor. AVAILABILITY: Mtreemix, a software package for estimating tree mixture models, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de. The raw cancer datasets and R code for the analysis with Cox models are available upon request from the corresponding author. Jörg Rahnenführer, Niko Beerenwinkel, Wolfgang A. Schulz, Christian Hartmann 0006, Andreas von Deimling, Bernd Wullich, Thomas Lengauer |
Bioinform. | 1 |
| 2005 | Confirmation of human protein interaction data by human expression dataabstractBACKGROUND: With microarray technology the expression of thousands of genes can be measured simultaneously. It is well known that the expression levels of genes of interacting proteins are correlated significantly more strongly in Saccharomyces cerevisiae than those of proteins that are not interacting. The objective of this work is to investigate whether this observation extends to the human genome. RESULTS: We investigated the quantitative relationship between expression levels of genes encoding interacting proteins and genes encoding random protein pairs. Therefore we studied 1369 interacting human protein pairs and human gene expression levels of 155 arrays. We were able to establish a statistically significantly higher correlation between the expression levels of genes whose proteins interact compared to random protein pairs. Additionally we were able to provide evidence that genes encoding proteins belonging to the same GO-class show correlated expression levels. CONCLUSION: This finding is concurrent with the naive hypothesis that the scales of production of interacting proteins are linked because an efficient interaction demands that involved proteins are available to some degree. The goal of further research in this field will be to understand the biological mechanisms behind this observation. Jörg Rahnenführer, Priti Talwar, Thomas Lengauer |
BMC Bioinform. | 2 |
| 2004 | Learning multiple evolutionary pathways from cross-sectional dataabstractWe introduce a mixture model of trees to describe evolutionary processes that are characterized by the accumulation of permanent genetic changes. The basic building block of the model is a directed weighted tree that generates a probability distribution on the set of all patterns of genetic events. We present an EM-like algorithm for learning a mixture model of K trees and show how to determine K with a maximum likelihood approach. As a case study we consider the accumulation of mutations in the HIV-1 reverse transcriptase that are associated with drug resistance. The fitted model is statistically validated as a density estimator and the stability of the model topology is analyzed. We obtain a generative probabilistic model for the development of drug resistance in HIV that agrees with biological knowledge. Further applications and extensions of the model are discussed. Niko Beerenwinkel, Jörg Rahnenführer, Martin Däumer, Daniel Hoffmann, Rolf Kaiser, Joachim Selbig, Thomas Lengauer |
RECOMB | 2 |
| 2004 | Predicting protein structure classes from function predictionsabstractMOTIVATION: We introduce a new approach to using the information contained in sequence-to-function prediction data in order to recognize protein template classes, a critical step in predicting protein structure. The data on which our method is based comprise probabilities of functional categories; for given query sequences these probabilities are obtained by a neural net that has previously been trained on a variety of functionally important features. On a training set of sequences we assess the relevance of individual functional categories for identifying a given structural family. Using a combination of the most relevant categories, the likelihood of a query sequence to belong to a specific family can be estimated. RESULTS: The performance of the method is evaluated using cross-validation. For a fixed structural family and for every sequence, a score is calculated that measures the evidence for family membership. Even for structural families of small size, family members receive significantly higher scores. For some examples, we show that the relevant functional features identified by this method are biologically meaningful. The proposed approach can be used to improve existing sequence-to-structure prediction methods. AVAILABILITY: Matlab code is available on request from the authors. The data are available at http://www.mpisb.mpg.de/~sommer/Fun2Struc/ Ingolf Sommer, Jörg Rahnenführer, Francisco S. Domingues, Ulrik de Lichtenberg, Thomas Lengauer |
Bioinform. | 2 |
| 2004 | Hybrid clustering for microarray image analysis combining intensity and shape featuresabstractBACKGROUND: Image analysis is the first crucial step to obtain reliable results from microarray experiments. First, areas in the image belonging to single spots have to be identified. Then, those target areas have to be partitioned into foreground and background. Finally, two scalar values for the intensities have to be extracted. These goals have been tackled either by spot shape methods or intensity histogram methods, but it would be desirable to have hybrid algorithms which combine the advantages of both approaches. RESULTS: A new robust and adaptive histogram type method is pixel clustering, which has been successfully applied for detecting and quantifying microarray spots. This paper demonstrates how the spot shape can be effectively integrated in this approach. Based on the clustering results, a bivalence mask is constructed. It estimates the expected spot shape and is used to filter the data, improving the results of the cluster algorithm. The quality measure 'stability' is defined and evaluated on a real data set. The improved clustering method is compared with the established Spot software on a data set with replicates. CONCLUSION: The new method presents a successful hybrid microarray image analysis solution. It incorporates both shape and histogram features and is specifically adapted to deal with typical microarray image characteristics. As a consequence of the filtering step pixels are divided into three groups, namely foreground, background and deletions. This allows a separate treatment of artifacts and their elimination from the further analysis. Jörg Rahnenführer, Daniel Bozinov |
BMC Bioinform. | 1 |
| 2002 | Unsupervised technique for robust target separation and analysis of DNA microarray spots through adaptive pixel clusteringabstractMOTIVATION: Microarray images challenge existing analytical methods in many ways given that gene spots are often comprised of characteristic imperfections. Irregular contours, donut shapes, artifacts, and low or heterogeneous expression impair corresponding values for red and green intensities as well as their ratio R/G. New approaches are needed to ensure accurate data extraction from these images. RESULTS: Herein we introduce a novel method for intensity assessment of gene spots. The technique is based on clustering pixels of a target area into foreground and background. For this purpose we implemented two clustering algorithms derived from k-means and Partitioning Around Medoids (PAM), respectively. Results from the analysis of real gene spots indicate that our approach performs superior to other existing analytical methods. This is particularly true for spots generally considered as problematic due to imperfections or almost absent expression. Both PX(PAM) and PX(KMEANS) prove to be highly robust against various types of artifacts through adaptive partitioning, which more correctly assesses expression intensity values. AVAILABILITY: The implementation of this method is a combination of two complementary tools Extractiff (Java) and Pixclust (free statistical language R), which are available upon request from the authors. Daniel Bozinov, Jörg Rahnenführer |
Bioinform. | 2 |