Lars Kaderali

dblp:09/835 · DBLP profile ↗
← Back
20ranked-venue papers
3as first author
6since 2021 · last 2026
0000-0002-2359-2294ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 18 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Theory of computation · 1
YearPublicationVenuePosition
2026 KiMeKo: A Collaborative AI Platform for Medical Device Development
abstract
KiMeKo (KI-Med-Kollaborationsplattform) is a publically funded collaborative research project that develops a sustainable AI-Med ecosystem for AI-based medical device development. The project runs from July 2024 to December 2027 and joins seven Northern German research institutions. KiMeKo addresses the complete development trajectory, from concept and data acquisition to validation, regulatory evidence generation, and approval-oriented documentation. The project contributes a practical toolchain and platform capabilities for non-experts and experts, including structured innovation support, uncertaintyaware sensor-data fusion, hybrid expert-system modeling, and workflow-guided data acquisition and anonymization. This paper summarizes project objectives, expected outputs, relevance to IEEE COMPSAC 2026 themes, and current progress. In particular, KiMeKo aligns with Applied AI and Smart & Connected Health by combining AI engineering, privacy-conscious data processing, and regulation-aware medical software development.
Serge Autexier, Nihat Ay, Stefan Fischer 0001, Lars Kaderali, Thomas Kirste, Martin Leucker, Christoph Lüth, Thomas Martinetz, Philipp Rostalski, Alexander Schlaefer, Frank Ückert
COMPSAC4
2024 Transparent AI Models for Meningococcal Meningitis Diagnosis: Evaluating Interpretability and Performance Metrics
abstract
Meningococcal meningitis, a severe bacterial infection, demands precise diagnosis and immediate intervention. This paper explores the potential of Artificial Intelligence (AI) in improving diagnostic accuracy and clinical decision-making for this condition. We assess interpretability and reliability by analyzing the decision-making process of Machine Learning (ML) models tailored to identify Meningococcal meningitis among different etiologies of meningitis. Initially, we train and test various ML models, including logistic regression, K-nearest neighbors, support vector machine, decision tree, gradient boosting, AdaBoost, random forest, and LightGBM classifier, on a dataset comprising 460 cases of Meningococcal meningitis and 474 cases of other types of meningitis. Among these models, the gradient-boosting model exhibits superior performance metrics, including accuracy (0.88), precision (0.92), recall (0.83), AUROC (0.93), and f1-score (0.87). Subsequently, we employ Explainable AI (XAI) tools (ELI5 and LIME) to elucidate the importance of features in the ML models. Our results highlight key factors contributing to model success, such as Neisseria meningitidis identified through cerebrospinal fluid (CSF) culture and latex agglutination, gram-negative diplococci in CSF smear examination, white cell counts in CSF, and patient age. Local explanations reveal the presence of neutrophils in CSF as a characteristic feature of Meningococcal meningitis. LIME analysis indicates the significance of low lymphocyte percentages and elevated white blood cell counts in predicting this condition. These findings underscore the effectiveness of integrating global and local interpretability techniques, aligning with expert knowledge, and emphasizing the importance of transparent AI models in clinical decision-making processes.
Aya Messai, Ahlem Drif, Amel Ouyahia, Meriem Guechi, Mounira Rais, Lars Kaderali, Hocine Cherifi
IS6
2024 DiscovEpi: automated whole proteome MHC-I-epitope prediction and visualization
abstract
BACKGROUND: T cells that recognize these complexes with their T cell receptor are activated and ideally eliminate infected cells. Prediction of putative peptides binding to MHC class I (MHC-I) is crucial for understanding pathogen recognition in specific immune responses and for supporting drug and vaccine design. There are reliable databases for epitope prediction algorithms available however they primarily focus on the prediction of epitopes in single immunogenic proteins. RESULTS: We have developed the tool DiscovEpi to establish an interface between whole proteomes and epitope prediction. The tool allows the automated identification of all potential MHC-I-binding peptides within a proteome and calculates the epitope density and average binding score for every protein, a protein-centric approach. DiscovEpi provides a convenient interface between automated multiple sequence extraction by organism and cell compartment from the database UniProt for subsequent epitope prediction via NetMHCpan. Furthermore, it allows ranking of proteins by their predicted immunogenicity on the one hand and comparison of different proteomes on the other. By applying the tool, we predict a higher immunogenic potential of membrane-associated proteins of SARS-CoV-2 compared to those of influenza A based on the presented metrics epitope density and binding score. This could be confirmed visually by comparing the epitope maps of the influenza A strain and SARS-CoV-2. CONCLUSION: Automated prediction of whole proteomes and the subsequent visualization of the location of putative epitopes on sequence-level facilitate the search for putative immunogenic proteins or protein regions and support the study of adaptive immune responses and vaccine design.
Cedric Mahncke, Frieder Schmiedeke, Stefan Simm, Lars Kaderali, Barbara M. Bröker, Ulrike Seifert, Clemens Cammann
BMC Bioinform.4
2023 Mathematical modeling of plus-strand RNA virus replication to identify broad-spectrum antiviral treatment strategies
abstract
Plus-strand RNA viruses are the largest group of viruses. Many are human pathogens that inflict a socio-economic burden. Interestingly, plus-strand RNA viruses share remarkable similarities in their replication. A hallmark of plus-strand RNA viruses is the remodeling of intracellular membranes to establish replication organelles (so-called "replication factories"), which provide a protected environment for the replicase complex, consisting of the viral genome and proteins necessary for viral RNA synthesis. In the current study, we investigate pan-viral similarities and virus-specific differences in the life cycle of this highly relevant group of viruses. We first measured the kinetics of viral RNA, viral protein, and infectious virus particle production of hepatitis C virus (HCV), dengue virus (DENV), and coxsackievirus B3 (CVB3) in the immuno-compromised Huh7 cell line and thus without perturbations by an intrinsic immune response. Based on these measurements, we developed a detailed mathematical model of the replication of HCV, DENV, and CVB3 and showed that only small virus-specific changes in the model were necessary to describe the in vitro dynamics of the different viruses. Our model correctly predicted virus-specific mechanisms such as host cell translation shut off and different kinetics of replication organelles. Further, our model suggests that the ability to suppress or shut down host cell mRNA translation may be a key factor for in vitro replication efficiency, which may determine acute self-limited or chronic infection. We further analyzed potential broad-spectrum antiviral treatment options in silico and found that targeting viral RNA translation, such as polyprotein cleavage and viral RNA synthesis, may be the most promising drug targets for all plus-strand RNA viruses. Moreover, we found that targeting only the formation of replicase complexes did not stop the in vitro viral replication early in infection, while inhibiting intracellular trafficking processes may even lead to amplified viral growth.
Carolin Zitzmann, Christopher Dächert, Bianca Schmid, Hilde van der Schaar, Martijn van Hemert, Alan S. Perelson, Frank J. M. van Kuppeveld, Ralf Bartenschlager, Marco Binder, Lars Kaderali
PLoS Comput. Biol.10
2021 Computational strategies to combat COVID-19: useful tools to accelerate SARS-CoV-2 and coronavirus research
abstract
SARS-CoV-2 (severe acute respiratory syndrome coronavirus 2) is a novel virus of the family Coronaviridae. The virus causes the infectious disease COVID-19. The biology of coronaviruses has been studied for many years. However, bioinformatics tools designed explicitly for SARS-CoV-2 have only recently been developed as a rapid reaction to the need for fast detection, understanding and treatment of COVID-19. To control the ongoing COVID-19 pandemic, it is of utmost importance to get insight into the evolution and pathogenesis of the virus. In this review, we cover bioinformatics workflows and tools for the routine detection of SARS-CoV-2 infection, the reliable analysis of sequencing data, the tracking of the COVID-19 pandemic and evaluation of containment measures, the study of coronavirus evolution, the discovery of potential drug targets and development of therapeutic strategies. For each tool, we briefly describe its use case and how it advances research specifically for SARS-CoV-2. All tools are free to use and available online, either through web applications or public code repositories. Contact:[email protected].
Franziska Hufsky, Kevin Lamkiewicz, Alexandre Almeida, Abdel Aouacheria, Cecilia N. Arighi, Alex Bateman, Jan Baumbach, Niko Beerenwinkel, Christian Brandt, Marco Cacciabue, Sara Chuguransky, Oliver Drechsel, Robert D. Finn, Adrian Fritz, Stephan Fuchs, Georges Hattab, Anne-Christin Hauschild, Dominik Heider, Marie Hoffmann, Martin Hölzer, Stefan Hoops, Lars Kaderali, Ioanna Kalvari, Max von Kleist, Renó Kmiecinski, Denise Kühnert, Gorka Lasso, Pieter Libin, Markus List, Hannah F. Löchel, Maria Jesus Martin, Roman Martin, Julian O. Matschinske, Alice C. McHardy, Pedro Mendes 0001, Jaina Mistry, Vincent Navratil, Eric P. Nawrocki, Áine Niamh O'toole, Nancy Ontiveros-Palacios, Anton I. Petrov, Guillermo Rangel-Pineros, Nicole Redaschi, Susanne Reimering, Knut Reinert, Lorna J. Richardson, David L. Robertson, Sepideh Sadegh, Joshua B. Singer, Kristof Theys, Chris Upton, Marius Welzel, Lowri Williams, Manja Marz
Briefings Bioinform.22
2021 From heterogeneous healthcare data to disease-specific biomarker networks: A hierarchical Bayesian network approach
abstract
In this work, we introduce an entirely data-driven and automated approach to reveal disease-associated biomarker and risk factor networks from heterogeneous and high-dimensional healthcare data. Our workflow is based on Bayesian networks, which are a popular tool for analyzing the interplay of biomarkers. Usually, data require extensive manual preprocessing and dimension reduction to allow for effective learning of Bayesian networks. For heterogeneous data, this preprocessing is hard to automatize and typically requires domain-specific prior knowledge. We here combine Bayesian network learning with hierarchical variable clustering in order to detect groups of similar features and learn interactions between them entirely automated. We present an optimization algorithm for the adaptive refinement of such group Bayesian networks to account for a specific target variable, like a disease. The combination of Bayesian networks, clustering, and refinement yields low-dimensional but disease-specific interaction networks. These networks provide easily interpretable, yet accurate models of biomarker interdependencies. We test our method extensively on simulated data, as well as on data from the Study of Health in Pomerania (SHIP-TREND), and demonstrate its effectiveness using non-alcoholic fatty liver disease and hypertension as examples. We show that the group network models outperform available biomarker scores, while at the same time, they provide an easily interpretable interaction network.
Ann-Kristin Becker, Marcus Dörr, Stephan B. Felix, Fabian Frost, Hans Jörgen Grabe, Markus M. Lerch, Matthias Nauck, Uwe Völker, Henry Völzke, Lars Kaderali
PLoS Comput. Biol.10
2020 Host factor prioritization for pan-viral genetic perturbation screens using random intercept models and network propagation
abstract
Genetic perturbation screens using RNA interference (RNAi) have been conducted successfully to identify host factors that are essential for the life cycle of bacteria or viruses. So far, most published studies identified host factors primarily for single pathogens. Furthermore, often only a small subset of genes, e.g., genes encoding kinases, have been targeted. Identification of host factors on a pan-pathogen level, i.e., genes that are crucial for the replication of a diverse group of pathogens has received relatively little attention, despite the fact that such common host factors would be highly relevant, for instance, for devising broad-spectrum anti-pathogenic drugs. Here, we present a novel two-stage procedure for the identification of host factors involved in the replication of different viruses using a combination of random effects models and Markov random walks on a functional interaction network. We first infer candidate genes by jointly analyzing multiple perturbations screens while at the same time adjusting for high variance inherent in these screens. Subsequently the inferred estimates are spread across a network of functional interactions thereby allowing for the analysis of missing genes in the biological studies, smoothing the effect sizes of previously found host factors, and considering a priori pathway information defined over edges of the network. We applied the procedure to RNAi screening data of four different positive-sense single-stranded RNA viruses, Hepatitis C virus, Chikungunya virus, Dengue virus and Severe acute respiratory syndrome coronavirus, and detected novel host factors, including UBC, PLCG1, and DYRK1B, which are predicted to significantly impact the replication cycles of these viruses. We validated the detected host factors experimentally using pharmacological inhibition and an additional siRNA screen and found that some of the predicted host factors indeed influence the replication of these pathogens.
Simon Dirmeier, Christopher Dächert, Martijn van Hemert, Ali Tas, Natacha S. Ogando, Frank J. M. van Kuppeveld, Ralf Bartenschlager, Lars Kaderali, Marco Binder, Niko Beerenwinkel
PLoS Comput. Biol.8
2020 Mathematical modeling of hepatitis C RNA replication, exosome secretion and virus release
abstract
Hepatitis C virus (HCV) causes acute hepatitis C and can lead to life-threatening complications if it becomes chronic. The HCV genome is a single plus strand of RNA. Its intracellular replication is a spatiotemporally coordinated process of RNA translation upon cell infection, RNA synthesis within a replication compartment, and virus particle production. While HCV is mainly transmitted via mature infectious virus particles, it has also been suggested that HCV-infected cells can secrete HCV RNA carrying exosomes that can infect cells in a receptor independent manner. In order to gain insight into these two routes of transmission, we developed a series of intracellular HCV replication models that include HCV RNA secretion and/or virus assembly and release. Fitting our models to in vitro data, in which cells were infected with HCV, suggests that initially most secreted HCV RNA derives from intracellular cytosolic plus-strand RNA, but subsequently secreted HCV RNA derives equally from the cytoplasm and the replication compartments. Furthermore, our model fits to the data suggest that the rate of virus assembly and release is limited by host cell resources. Including the effects of direct acting antivirals in our models, we found that in spite of decreasing intracellular HCV RNA and extracellular virus concentration, low level HCV RNA secretion may continue as long as intracellular RNA is available. This may possibly explain the presence of detectable levels of plasma HCV RNA at the end of treatment even in patients that ultimately attain a sustained virologic response.
Carolin Zitzmann, Lars Kaderali, Alan S. Perelson
PLoS Comput. Biol.2
2015 lpNet: a linear programming approach to reconstruct signal transduction networks
abstract
UNLABELLED: With the widespread availability of high-throughput experimental technologies it has become possible to study hundreds to thousands of cellular factors simultaneously, such as coding- or non-coding mRNA or protein concentrations. Still, extracting information about the underlying regulatory or signaling interactions from these data remains a difficult challenge. We present a flexible approach towards network inference based on linear programming. Our method reconstructs the interactions of factors from a combination of perturbation/non-perturbation and steady-state/time-series data. We show both on simulated and real data that our methods are able to reconstruct the underlying networks fast and efficiently, thus shedding new light on biological processes and, in particular, into disease's mechanisms of action. We have implemented the approach as an R package available through bioconductor. AVAILABILITY AND IMPLEMENTATION: This R package is freely available under the Gnu Public License (GPL-3) from bioconductor.org (http://bioconductor.org/packages/release/bioc/html/lpNet.html) and is compatible with most operating systems (Windows, Linux, Mac OS) and hardware architectures. CONTACT: [email protected] SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Marta R. A. Matos, Bettina Knapp, Lars Kaderali
Bioinform.3
2014 Dynamic Probabilistic Threshold Networks to Infer Signaling Pathways from Time-Course Perturbation Data
abstract
BACKGROUND: Network inference deals with the reconstruction of molecular networks from experimental data. Given N molecular species, the challenge is to find the underlying network. Due to data limitations, this typically is an ill-posed problem, and requires the integration of prior biological knowledge or strong regularization. We here focus on the situation when time-resolved measurements of a system's response after systematic perturbations are available. RESULTS: We present a novel method to infer signaling networks from time-course perturbation data. We utilize dynamic Bayesian networks with probabilistic Boolean threshold functions to describe protein activation. The model posterior distribution is analyzed using evolutionary MCMC sampling and subsequent clustering, resulting in probability distributions over alternative networks. We evaluate our method on simulated data, and study its performance with respect to data set size and levels of noise. We then use our method to study EGF-mediated signaling in the ERBB pathway. CONCLUSIONS: Dynamic Probabilistic Threshold Networks is a new method to infer signaling networks from time-series perturbation data. It exploits the dynamic response of a system after external perturbation for network reconstruction. On simulated data, we show that the approach outperforms current state of the art methods. On the ERBB data, our approach recovers a significant fraction of the known interactions, and predicts novel mechanisms in the ERBB pathway.
Narsis A. Kiani, Lars Kaderali
BMC Bioinform.2
2013 Mining Quasi-Bicliques from HIV-1-Human Protein Interaction Network: A Multiobjective Biclustering Approach
abstract
In this work, we model the problem of mining quasi-bicliques from weighted viral-host protein-protein interaction network as a biclustering problem for identifying strong interaction modules. In this regard, a multiobjective genetic algorithm-based biclustering technique is proposed that simultaneously optimizes three objective functions to obtain dense biclusters having high mean interaction strengths. The performance of the proposed technique has been compared with that of other existing biclustering methods on an artificial data. Subsequently, the proposed biclustering method is applied on the records of biologically validated and predicted interactions between a set of HIV-1 proteins and a set of human proteins to identify strong interaction modules. For this, the entire interaction information is realized as a bipartite graph. We have further investigated the biological significance of the obtained biclusters. The human proteins involved in the strong interaction module have been found to share common biological properties and they are identified as the gateways of viral infection leading to various diseases. These human proteins can be potential drug targets for developing anti-HIV drugs.
Ujjwal Maulik, Anirban Mukhopadhyay 0001, Malay Bhattacharyya 0001, Lars Kaderali, Benedikt Brors, Sanghamitra Bandyopadhyay, Roland Eils
IEEE ACM Trans. Comput. Biol. Bioinform.4
2011 Normalizing for individual cell population context in the analysis of high-content cellular screens
abstract
BACKGROUND: High-content, high-throughput RNA interference (RNAi) offers unprecedented possibilities to elucidate gene function and involvement in biological processes. Microscopy based screening allows phenotypic observations at the level of individual cells. It was recently shown that a cell's population context significantly influences results. However, standard analysis methods for cellular screens do not currently take individual cell data into account unless this is important for the phenotype of interest, i.e. when studying cell morphology. RESULTS: We present a method that normalizes and statistically scores microscopy based RNAi screens, exploiting individual cell information of hundreds of cells per knockdown. Each cell's individual population context is employed in normalization. We present results on two infection screens for hepatitis C and dengue virus, both showing considerable effects on observed phenotypes due to population context. In addition, we show on a non-virus screen that these effects can be found also in RNAi data in the absence of any virus. Using our approach to normalize against these effects we achieve improved performance in comparison to an analysis without this normalization and hit scoring strategy. Furthermore, our approach results in the identification of considerably more significantly enriched pathways in hepatitis C virus replication than using a standard analysis approach. CONCLUSIONS: Using a cell-based analysis and normalization for population context, we achieve improved sensitivity and specificity not only on a individual protein level, but especially also on a pathway level. This leads to the identification of new host dependency factors of the hepatitis C and dengue viruses and higher reproducibility of results.
Bettina Knapp, Ilka Rebhan, Anil Kumar 0006, Petr Matula, Narsis A. Kiani, Marco Binder, Holger Erfle, Karl Rohr, Roland Eils, Ralf Bartenschlager, Lars Kaderali
BMC Bioinform.11
2010 Detecting host factors involved in virus infection by observing the clustering of infected cells in siRNA screening images
abstract
MOTIVATION: Detecting human proteins that are involved in virus entry and replication is facilitated by modern high-throughput RNAi screening technology. However, hit lists from different laboratories have shown only little consistency. This may be caused by not only experimental discrepancies, but also not fully explored possibilities of the data analysis. We wanted to improve reliability of such screens by combining a population analysis of infected cells with an established dye intensity readout. RESULTS: Viral infection is mainly spread by cell-cell contacts and clustering of infected cells can be observed during spreading of the infection in situ and in vivo. We employed this clustering feature to define knockdowns which harm viral infection efficiency of human Hepatitis C Virus. Images of knocked down cells for 719 human kinase genes were analyzed with an established point pattern analysis method (Ripley's K-function) to detect knockdowns in which virally infected cells did not show any clustering and therefore were hindered to spread their infection to their neighboring cells. The results were compared with a statistical analysis using a common intensity readout of the GFP-expressing viruses and a luciferase-based secondary screen yielding five promising host factors which may suit as potential targets for drug therapy. CONCLUSION: We report of an alternative method for high-throughput imaging methods to detect host factors being relevant for the infection efficiency of viruses. The method is generic and has the potential to be used for a large variety of different viruses and treatments being screened by imaging techniques.
Apichat Suratanee, Ilka Rebhan, Petr Matula, Anil Kumar 0006, Lars Kaderali, Karl Rohr, Ralf Bartenschlager, Roland Eils, Rainer König
Bioinform.5
2009 Reconstructing signaling pathways from RNAi data using probabilistic Boolean threshold networks
abstract
MOTIVATION: The reconstruction of signaling pathways from gene knockdown data is a novel research field enabled by developments in RNAi screening technology. However, while RNA interference is a powerful technique to identify genes related to a phenotype of interest, their placement in the corresponding pathways remains a challenging problem. Difficulties are aggravated if not all pathway components can be observed after each knockdown, but readouts are only available for a small subset. We are then facing the problem of reconstructing a network from incomplete data. RESULTS: We infer pathway topologies from gene knockdown data using Bayesian networks with probabilistic Boolean threshold functions. To deal with the problem of underdetermined network parameters, we employ a Bayesian learning approach, in which we can integrate arbitrary prior information on the network under consideration. Missing observations are integrated out. We compute the exact likelihood function for smaller networks, and use an approximation to evaluate the likelihood for larger networks. The posterior distribution is evaluated using mode hopping Markov chain Monte Carlo. Distributions over topologies and parameters can then be used to design additional experiments. We evaluate our approach on a small artificial dataset, and present inference results on RNAi data from the Jak/Stat pathway in a human hepatoma cell line.
Lars Kaderali, Eva Dazert, Ulf Zeuge, Michael Frese, Ralf Bartenschlager
Bioinform.1
2009 RNAither, an automated pipeline for the statistical analysis of high-throughput RNAi screens
abstract
SUMMARY: We present RNAither, a package for the free statistical environment R which performs an analysis of high-throughput RNA interference (RNAi) knock-down experiments, generating lists of relevant genes and pathways out of raw experimental data. The library provides a quality assessment of the signal intensities, as well as a broad range of options for data normalization, different statistical tests for the identification of significant siRNAs, and a significance analysis of the biological processes involving corresponding genes. The results of the analysis are presented as a set of HTML pages. Additionally, all values and plots are available as either text files or pdf and png files. AVAILABILITY: http://bioconductor.org/
Nora Rieber, Bettina Knapp, Roland Eils, Lars Kaderali
Bioinform.4
2009 Reconstructing nonlinear dynamic models of gene regulation using stochastic sampling
abstract
BACKGROUND: The reconstruction of gene regulatory networks from time series gene expression data is one of the most difficult problems in systems biology. This is due to several reasons, among them the combinatorial explosion of possible network topologies, limited information content of the experimental data with high levels of noise, and the complexity of gene regulation at the transcriptional, translational and post-translational levels. At the same time, quantitative, dynamic models, ideally with probability distributions over model topologies and parameters, are highly desirable. RESULTS: We present a novel approach to infer such models from data, based on nonlinear differential equations, which we embed into a stochastic Bayesian framework. We thus address both the stochasticity of experimental data and the need for quantitative dynamic models. Furthermore, the Bayesian framework allows it to easily integrate prior knowledge into the inference process. Using stochastic sampling from the Bayes' posterior distribution, our approach can infer different likely network topologies and model parameters along with their respective probabilities from given data. We evaluate our approach on simulated data and the challenge #3 data from the DREAM 2 initiative. On the simulated data, we study effects of different levels of noise and dataset sizes. Results on real data show that the dynamics and main regulatory interactions are correctly reconstructed. CONCLUSIONS: Our approach combines dynamic modeling using differential equations with a stochastic learning framework, thus bridging the gap between biophysical modeling and stochastic inference approaches. Results show that the method can reap the advantages of both worlds, and allows the reconstruction of biophysically accurate dynamic models from noisy data. In addition, the stochastic learning framework used permits the computation of probability distributions over models and model parameters, which holds interesting prospects for experimental design purposes.
Johanna Mazur, Daniel Ritter 0001, Gerhard Reinelt, Lars Kaderali
BMC Bioinform.4
2009 Inference of an oscillating model for the yeast cell cycle
Nicole Radde, Lars Kaderali
Discret. Appl. Math.2
2006 CASPAR: a hierarchical bayesian approach to predict survival times in cancer from gene expression data
abstract
MOTIVATION: DNA microarrays allow the simultaneous measurement of thousands of gene expression levels in any given patient sample. Gene expression data have been shown to correlate with survival in several cancers, however, analysis of the data is difficult, since typically at most a few hundred patients are available, resulting in severely underdetermined regression or classification models. Several approaches exist to classify patients in different risk classes, however, relatively little has been done with respect to the prediction of actual survival times. We introduce CASPAR, a novel method to predict true survival times for the individual patient based on microarray measurements. CASPAR is based on a multivariate Cox regression model that is embedded in a Bayesian framework. A hierarchical prior distribution on the regression parameters is specifically designed to deal with high dimensionality (large number of genes) and low sample size settings, that are typical for microarray measurements. This enables CASPAR to automatically select small, most informative subsets of genes for prediction. RESULTS: Validity of the method is demonstrated on two publicly available datasets on diffuse large B-cell lymphoma (DLBCL) and on adenocarcinoma of the lung. The method successfully identifies long and short survivors, with high sensitivity and specificity. We compare our method with two alternative methods from the literature, demonstrating superior results of our approach. In addition, we show that CASPAR can further refine predictions made using clinical scoring systems such as the International Prognostic Index (IPI) for DLBCL and clinical staging for lung cancer, thus providing an additional tool for the clinician. An analysis of the genes identified confirms previously published results, and furthermore, new candidate genes correlated with survival are identified.
Lars Kaderali, Thomas Zander, Ulrich Faigle, Joachim L. Schultze, Rainer Schrader
Bioinform.1
2005 A fractional programming approach to efficient DNA melting temperature calculation
abstract
MOTIVATION: In a wide range of experimental techniques in biology, there is a need for an efficient method to calculate the melting temperature of pairings of two single DNA strands. Avoiding cross-hybridization when choosing primers for the polymerase chain reaction or selecting probes for large-scale DNA assays are examples where the exact determination of melting temperatures is important. Beyond being exact, the method has to be efficient, as these techniques often require the simultaneous calculation of melting temperatures of up to millions of possible pairings. The problem is to simultaneously determine the most stable alignment of two sequences, including potential loops and bulges, and calculate the corresponding melting temperature. RESULTS: As the melting temperature can be expressed as a fraction in terms of enthalpy and entropy differences of the corresponding annealing reaction, we propose to use a fractional programming algorithm, the Dinkelbach algorithm, to solve the problem. To calculate the required differences of enthalpy and entropy, the Nearest Neighbor model is applied. Using this model, the substeps of the Dinkelbach algorithm in our problem setting turn out to be calculations of alignments which optimize an additive score function. Thus, the usual dynamic programming techniques can be applied. The result is an efficient algorithm to determine melting temperatures of two DNA strands, suitable for large-scale applications such as primer or probe design. AVAILABILITY: The software is available for academic purposes from the authors. A web interface is provided at http://www.zaik.uni-koeln.de/bioinformatik/fptm.html
Markus Leber, Lars Kaderali, Alexander Schönhuth, Rainer Schrader
Bioinform.2
2002 Selecting signature oligonucleotides to identify organisms using DNA arrays
abstract
MOTIVATION: DNA arrays are a very useful tool to quickly identify biological agents present in some given sample, e.g. to identify viruses causing disease, for quality control in the food industry, or to determine bacteria contaminating drinking water. The selection of specific oligos to attach to the array surface is a relevant problem in the experiment design process. Given a set S of genomic sequences (the target sequences), the task is to find at least one oligonucleotide, called probe, for each sequence in S. This probe will be attached to the array surface, and must be chosen in a way that it will not hybridize to any other sequence but the intended target. Furthermore, all probes on the array must hybridize to their intended targets under the same reaction conditions, most importantly at the temperature T at which the experiment is conducted. RESULTS: We present an efficient algorithm for the probe design problem. Melting temperatures are calculated for all possible probe-target interactions using an extended nearest-neighbor model, allowing for both non-Watson-Crick base-pairing and unpaired bases within a duplex. To compute temperatures efficiently, a combination of suffix trees and dynamic programming based alignment algorithms is introduced. Additional filtering steps during preprocessing increase the speed of the computation. The practicability of the algorithms is demonstrated by two case studies: The identification of HIV-1 subtypes, and of 28S rDNA sequences from >or=400 organisms.
Lars Kaderali, Alexander Schliep
Bioinform.1