EDBT 2026 Demo / reviewers in the wild / expert
Joachim Selbig
dblp:00/6376
· DBLP profile ↗
27ranked-venue papers
1as first author
0since 2021 · last 2015
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 25 · 1 first-authorArtificial intelligence and machine learning · 1Software engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1Theory of computation · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
16 papers |
Bioinformatics and computational biology · 99% Computational science and engineering · 1% | |
| Theoretical computer science
2 papers |
Algorithms and data structures · 100% |
Topics — the 26 heaviest of 30, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Bioinformatics and computational biology › systems biology
metabolic network analysis |
0.5 | 4 | 2015 | Stoichiometric capacitance reveals the theoretical capabilities of metabolic networks · Bioinform. 2012 A MATLAB toolbox for structural kinetic modeling · Bioinform. 2012 Mass-balanced randomization of metabolic networks · Bioinform. 2011 |
Bioinformatics and computational biology
systems biology |
0.3 | 2 | 2012 | Stoichiometric capacitance reveals the theoretical capabilities of metabolic networks · Bioinform. 2012 A MATLAB toolbox for structural kinetic modeling · Bioinform. 2012 |
Bioinformatics and computational biology › network bioinformatics › biological network analysis › network analysis
network stability |
0.2 | 1 | 2015 | Refined elasticity sampling for Monte Carlo-based identification of stabilizing network patterns · Bioinform. 2015 |
Bioinformatics and computational biology
metabolomics |
0.1 | 3 | 2005 | Non-linear PCA: a missing data approach · Bioinform. 2005 Metabolite fingerprinting: detecting biological features by independent component analysis · Bioinform. 2004 Threshold extraction in metabolite concentration data · Bioinform. 2004 |
Bioinformatics and computational biology
network randomization |
0.1 | 1 | 2011 | Mass-balanced randomization of metabolic networks · Bioinform. 2011 |
Algorithms and data structures
polynomial-time algorithms |
0.1 | 1 | 2011 | Mass-balanced randomization of metabolic networks · Bioinform. 2011 |
Bioinformatics and computational biology › systems bioinformatics › pathway analysis
metabolic pathway analysis |
0.1 | 1 | 2007 | From structure to dynamics of metabolic pathways: application to the plant mitochondrial TCA cycle · Bioinform. 2007 |
Bioinformatics and computational biology
omics data analysis |
0.1 | 1 | 2007 | pcaMethods - a bioconductor package providing PCA methods for incomplete data · Bioinform. 2007 |
Bioinformatics and computational biology
gene expression analysis |
0.1 | 2 | 2004 | Hypothesis-driven approach to predict transcriptional units from gene expression data · Bioinform. 2004 MetaGeneAlyse: analysis of integrated transcriptional and metabolite data · Bioinform. 2003 |
Bioinformatics and computational biology › computational microbiology
drug resistance prediction |
0.1 | 1 | 2005 | Computational methods for the design of effective therapies against drug resistant HIV strains · Bioinform. 2005 |
Bioinformatics and computational biology › genomics
viral genomics |
0.1 | 1 | 2005 | Computational methods for the design of effective therapies against drug resistant HIV strains · Bioinform. 2005 |
Bioinformatics and computational biology › clinical bioinformatics
HIV drug resistance |
0.0 | 1 | 2004 | Learning multiple evolutionary pathways from cross-sectional data · RECOMB 2004 |
Bioinformatics and computational biology › metabolomics
metabolite quantification |
0.0 | 1 | 2004 | Threshold extraction in metabolite concentration data · Bioinform. 2004 |
Bioinformatics and computational biology › genome annotation
transcriptional unit prediction |
0.0 | 1 | 2004 | Hypothesis-driven approach to predict transcriptional units from gene expression data · Bioinform. 2004 |
Bioinformatics and computational biology › systems biology
flux variability analysis |
0.0 | 1 | 2012 | Stoichiometric capacitance reveals the theoretical capabilities of metabolic networks · Bioinform. 2012 |
Bioinformatics and computational biology › multi-omics data integration
integrative omics analysis |
0.0 | 1 | 2003 | MetaGeneAlyse: analysis of integrated transcriptional and metabolite data · Bioinform. 2003 |
Bioinformatics and computational biology › protein structure prediction
consensus prediction |
0.0 | 1 | 1999 | Decision tree-based formation of consensus protein secondary structure prediction · Bioinform. 1999 |
Bioinformatics and computational biology
protein structure prediction |
0.0 | 1 | 1999 | Decision tree-based formation of consensus protein secondary structure prediction · Bioinform. 1999 |
Bioinformatics and computational biology › protein structure prediction
secondary structure prediction |
0.0 | 1 | 1999 | Decision tree-based formation of consensus protein secondary structure prediction · Bioinform. 1999 |
Bioinformatics and computational biology › systems biology › metabolic network analysis
elementary flux modes |
0.0 | 1 | 2007 | From structure to dynamics of metabolic pathways: application to the plant mitochondrial TCA cycle · Bioinform. 2007 |
Computational science and engineering › high-dimensional data analysis
nonlinear dimensionality reduction |
0.0 | 1 | 2005 | Non-linear PCA: a missing data approach · Bioinform. 2005 |
Bioinformatics and computational biology
sequence analysis |
0.0 | 1 | 2005 | Computational methods for the design of effective therapies against drug resistant HIV strains · Bioinform. 2005 |
Bioinformatics and computational biology › genome annotation
operon prediction |
0.0 | 1 | 2004 | Hypothesis-driven approach to predict transcriptional units from gene expression data · Bioinform. 2004 |
Bioinformatics and computational biology › biological database
pathway database |
0.0 | 1 | 2004 | PaVESy: Pathway Visualization and Editing System · Bioinform. 2004 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
concept learning |
0.0 | 1 | 1981 | Concept Learning by Structured Examples - An Algebraic Approach · IJCAI 1981 |
Algorithms and data structures › symbolic computation › computational algebra
algebraic algorithms |
0.0 | 1 | 1981 | Concept Learning by Structured Examples - An Algebraic Approach · IJCAI 1981 |
Methods — techniques the papers use, named apart from their topics
monte carlo sampling · 0.4switch randomization · 0.2null model sampling · 0.2jacobian matrix analysis · 0.2stoichiometric capacitance · 0.1parameter assignment · 0.1constraint-based modeling · 0.1principal component analysis · 0.1parametric differential equation modeling · 0.1missing value imputation · 0.1algebraic approach · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2015 | Refined elasticity sampling for Monte Carlo-based identification of stabilizing network patternsabstractMOTIVATION: Structural kinetic modelling (SKM) is a framework to analyse whether a metabolic steady state remains stable under perturbation, without requiring detailed knowledge about individual rate equations. It provides a representation of the system's Jacobian matrix that depends solely on the network structure, steady state measurements, and the elasticities at the steady state. For a measured steady state, stability criteria can be derived by generating a large number of SKMs with randomly sampled elasticities and evaluating the resulting Jacobian matrices. The elasticity space can be analysed statistically in order to detect network positions that contribute significantly to the perturbation response. Here, we extend this approach by examining the kinetic feasibility of the elasticity combinations created during Monte Carlo sampling. RESULTS: Using a set of small example systems, we show that the majority of sampled SKMs would yield negative kinetic parameters if they were translated back into kinetic models. To overcome this problem, a simple criterion is formulated that mitigates such infeasible models. After evaluating the small example pathways, the methodology was used to study two steady states of the neuronal TCA cycle and the intrinsic mechanisms responsible for their stability or instability. The findings of the statistical elasticity analysis confirm that several elasticities are jointly coordinated to control stability and that the main source for potential instabilities are mutations in the enzyme alpha-ketoglutarate dehydrogenase. Dorothee Childs, Sergio Grimbs, Joachim Selbig |
Bioinform. | 3 |
| 2012 | A MATLAB toolbox for structural kinetic modelingabstractSUMMARY: Structural kinetic modeling (SKM) enables the analysis of dynamical properties of metabolic networks solely based on topological information and experimental data. Current SKM-based experiments are hampered by the time-intensive process of assigning model parameters and choosing appropriate sampling intervals for Monte-Carlo experiments. We introduce a toolbox for the automatic and efficient construction and evaluation of structural kinetic models (SK models). Quantitative and qualitative analyses of network stability properties are performed in an automated manner. We illustrate the model building and analysis process in detailed example scripts that provide toolbox implementations of previously published literature models. AVAILABILITY: The source code is freely available for download at http://bioinformatics.uni-potsdam.de/projects/skm. CONTACT: [email protected]. Dorothee Girbig, Joachim Selbig, Sergio Grimbs |
Bioinform. | 2 |
| 2012 | Stoichiometric capacitance reveals the theoretical capabilities of metabolic networksabstractMOTIVATION: Metabolic engineering aims at modulating the capabilities of metabolic networks by changing the activity of biochemical reactions. The existing constraint-based approaches for metabolic engineering have proven useful, but are limited only to reactions catalogued in various pathway databases. RESULTS: We consider the alternative of designing synthetic strategies which can be used not only to characterize the maximum theoretically possible product yield but also to engineer networks with optimal conversion capability by using a suitable biochemically feasible reaction called 'stoichiometric capacitance'. In addition, we provide a theoretical solution for decomposing a given stoichiometric capacitance over a set of known enzymatic reactions. We determine the stoichiometric capacitance for genome-scale metabolic networks of 10 organisms from different kingdoms of life and examine its implications for the alterations in flux variability patterns. Our empirical findings suggest that the theoretical capacity of metabolic networks comes at a cost of dramatic system's changes. CONTACT: [email protected], or [email protected] SUPPLEMENTARY INFORMATION: Supplementary tables are available at Bioinformatics online. Abdelhalim Larhlimi, Georg Basler, Sergio Grimbs, Joachim Selbig, Zoran Nikoloski |
Bioinform. | 4 |
| 2012 | F2C2: a fast tool for the computation of flux coupling in genome-scale metabolic networksabstractBACKGROUND: Flux coupling analysis (FCA) has become a useful tool in the constraint-based analysis of genome-scale metabolic networks. FCA allows detecting dependencies between reaction fluxes of metabolic networks at steady-state. On the one hand, this can help in the curation of reconstructed metabolic networks by verifying whether the coupling between reactions is in agreement with the experimental findings. On the other hand, FCA can aid in defining intervention strategies to knock out target reactions. RESULTS: We present a new method F2C2 for FCA, which is orders of magnitude faster than previous approaches. As a consequence, FCA of genome-scale metabolic networks can now be performed in a routine manner. CONCLUSIONS: We propose F2C2 as a fast tool for the computation of flux coupling in genome-scale metabolic networks. F2C2 is freely available for non-commercial use at https://sourceforge.net/projects/f2c2/files/. Abdelhalim Larhlimi, László Dávid, Joachim Selbig, Alexander Bockmayr |
BMC Bioinform. | 3 |
| 2011 | Mass-balanced randomization of metabolic networksabstractMOTIVATION: Network-centered studies in systems biology attempt to integrate the topological properties of biological networks with experimental data in order to make predictions and posit hypotheses. For any topology-based prediction, it is necessary to first assess the significance of the analyzed property in a biologically meaningful context. Therefore, devising network null models, carefully tailored to the topological and biochemical constraints imposed on the network, remains an important computational problem. RESULTS: We first review the shortcomings of the existing generic sampling scheme-switch randomization-and explain its unsuitability for application to metabolic networks. We then devise a novel polynomial-time algorithm for randomizing metabolic networks under the (bio)chemical constraint of mass balance. The tractability of our method follows from the concept of mass equivalence classes, defined on the representation of compounds in the vector space over chemical elements. We finally demonstrate the uniformity of the proposed method on seven genome-scale metabolic networks, and empirically validate the theoretical findings. The proposed method allows a biologically meaningful estimation of significance for metabolic network properties. Georg Basler, Oliver Ebenhöh, Joachim Selbig, Zoran Nikoloski |
Bioinform. | 3 |
| 2010 | Biomarker discovery in heterogeneous tissue samples -taking the in-silico deconfounding approachabstractBACKGROUND: For heterogeneous tissues, such as blood, measurements of gene expression are confounded by relative proportions of cell types involved. Conclusions have to rely on estimation of gene expression signals for homogeneous cell populations, e.g. by applying micro-dissection, fluorescence activated cell sorting, or in-silico deconfounding. We studied feasibility and validity of a non-negative matrix decomposition algorithm using experimental gene expression data for blood and sorted cells from the same donor samples. Our objective was to optimize the algorithm regarding detection of differentially expressed genes and to enable its use for classification in the difficult scenario of reversely regulated genes. This would be of importance for the identification of candidate biomarkers in heterogeneous tissues. RESULTS: Experimental data and simulation studies involving noise parameters estimated from these data revealed that for valid detection of differential gene expression, quantile normalization and use of non-log data are optimal. We demonstrate the feasibility of predicting proportions of constituting cell types from gene expression data of single samples, as a prerequisite for a deconfounding-based classification approach.Classification cross-validation errors with and without using deconfounding results are reported as well as sample-size dependencies. Implementation of the algorithm, simulation and analysis scripts are available. CONCLUSIONS: The deconfounding algorithm without decorrelation using quantile normalization on non-log data is proposed for biomarkers that are difficult to detect, and for cases where confounding by varying proportions of cell types is the suspected reason. In this case, a deconfounding ranking approach can be used as a powerful alternative to, or complement of, other statistical learning approaches to define candidate biomarkers for molecular diagnosis and prediction in biomedicine, in realistically noisy conditions and with moderate sample sizes. Dirk Repsilber, Sabine Kern, Anna Telaar, Gerhard Walzl, Gillian F. Black, Joachim Selbig, Shreemanta K. Parida, Stefan H. E. Kaufmann, Marc Jacobsen |
BMC Bioinform. | 6 |
| 2009 | Detection and characterization of 3D-signature phosphorylation site motifs and their contribution towards improved phosphorylation site prediction in proteinsabstractBACKGROUND: Phosphorylation of proteins plays a crucial role in the regulation and activation of metabolic and signaling pathways and constitutes an important target for pharmaceutical intervention. Central to the phosphorylation process is the recognition of specific target sites by protein kinases followed by the covalent attachment of phosphate groups to the amino acids serine, threonine, or tyrosine. The experimental identification as well as computational prediction of phosphorylation sites (P-sites) has proved to be a challenging problem. Computational methods have focused primarily on extracting predictive features from the local, one-dimensional sequence information surrounding phosphorylation sites. RESULTS: We characterized the spatial context of phosphorylation sites and assessed its usability for improved phosphorylation site predictions. We identified 750 non-redundant, experimentally verified sites with three-dimensional (3D) structural information available in the protein data bank (PDB) and grouped them according to their respective kinase family. We studied the spatial distribution of amino acids around phosphorserines, phosphothreonines, and phosphotyrosines to extract signature 3D-profiles. Characteristic spatial distributions of amino acid residue types around phosphorylation sites were indeed discernable, especially when kinase-family-specific target sites were analyzed. To test the added value of using spatial information for the computational prediction of phosphorylation sites, Support Vector Machines were applied using both sequence as well as structural information. When compared to sequence-only based prediction methods, a small but consistent performance improvement was obtained when the prediction was informed by 3D-context information. CONCLUSION: While local one-dimensional amino acid sequence information was observed to harbor most of the discriminatory power, spatial context information was identified as relevant for the recognition of kinases and their cognate target sites and can be used for an improved prediction of phosphorylation sites. A web-based service (Phos3D) implementing the developed structure-based P-site prediction method has been made available at (http://phos3d.mpimp-golm.mpg.de). Pawel Durek, Christian Schudoma, Wolfram Weckwerth, Joachim Selbig, Dirk Walther 0001 |
BMC Bioinform. | 4 |
| 2008 | Hardness and Approximability of the Inverse Scope Problem
Zoran Nikoloski, Sergio Grimbs, Joachim Selbig, Oliver Ebenhöh |
WABI | 3 |
| 2007 | pcaMethods - a bioconductor package providing PCA methods for incomplete dataabstractAbstract Summary: pcaMethods is a Bioconductor compliant library for computing principal component analysis (PCA) on incomplete data sets. The results can be analyzed directly or used to estimate missing values to enable the use of missing value sensitive statistical methods. The package was mainly developed with microarray and metabolite data sets in mind, but can be applied to any other incomplete data set as well. Availability: http://www.bioconductor.org Contact: [email protected] Supplementary information: Please visit our webpage at http://bioinformatics.mpimp-golm.mpg.de/ Wolfram Stacklies, Henning Redestig, Matthias Scholz, Dirk Walther 0001, Joachim Selbig |
Bioinform. | 5 |
| 2007 | From structure to dynamics of metabolic pathways: application to the plant mitochondrial TCA cycleabstractMOTIVATION: Mitochondrial metabolism, dominated by the reactions of the tricarboxylic acid (TCA) cycle, is of vital importance for a wide range of metabolic processes. In particular for autotrophic tissue, such as plant leaves, the TCA cycle marks the point of divergence of anabolic pathways and plays an essential role in biosynthesis. However, despite extensive knowledge about its stoichiometric properties, the function and the dynamical capabilities of the TCA cycle remain largely unknown. METHODS AND RESULTS: Based on a recently proposed formalism, we investigate the dynamic and functional properties of the mitochondrial TCA cycle of plants. Starting with the structural properties, as described by the elementary flux modes of the system, we aim for the transition from structure to the dynamics of the TCA cycle. Using a parametric description of the system, encompassing all possible differential equations and parameter values, we detect and quantify regimes of different dynamic behavior. Optimizing the system with respect to dynamic stability, we demonstrate that maximal stability is associated with specific (relative) metabolite concentrations and flux values that are subsequently compared to the experimental literature. Our analysis also serves as a general example how to elucidate the transition from the structure to the dynamics of metabolic pathways. Ralf Steuer, Adriano Nunes Nesi, Alisdair R. Fernie, Thilo Gross, Bernd Blasius, Joachim Selbig |
Bioinform. | 6 |
| 2007 | ProMEX: a mass spectral reference database for proteins and protein phosphorylation sitesabstractBACKGROUND: In the last decade, techniques were established for the large scale genome-wide analysis of proteins, RNA, and metabolites, and database solutions have been developed to manage the generated data sets. The Golm Metabolome Database for metabolite data (GMD) represents one such effort to make these data broadly available and to interconnect the different molecular levels of a biological system 1. As data interpretation in the light of already existing data becomes increasingly important, these initiatives are an essential part of current and future systems biology. RESULTS: A mass spectral library consisting of experimentally derived tryptic peptide product ion spectra was generated based on liquid chromatography coupled to ion trap mass spectrometry (LC-IT-MS). Protein samples derived from Arabidopsis thaliana, Chlamydomonas reinhardii, Medicago truncatula, and Sinorhizobium meliloti were analysed. With currently 4,557 manually validated spectra associated with 4,226 unique peptides from 1,367 proteins, the database serves as a continuously growing reference data set and can be used for protein identification and quantification in uncharacterized biological samples. For peptide identification, several algorithms were implemented based on a recently published study for peptide mass fingerprinting 2 and tested for false positive and negative rates. An algorithm which considers intensity distribution for match correlation scores was found to yield best results. For proof of concept, an LC-IT-MS analysis of a tryptic leaf protein digest was converted to mzData format and searched against the mass spectral library. The utility of the mass spectral library was also tested for the identification of phosphorylated tryptic peptides. We included in vivo phosphorylation sites of Arabidopsis thaliana proteins and the identification performance was found to be improved compared to genome-based search algorithms. Protein identification by ProMEX is linked to other levels of biological organization such as metabolite, pathway, and transcript data. The database is further connected to annotation and classification services via BioMoby. CONCLUSION: The ProMEX protein/peptide database represents a mass spectral reference library with the capability of matching unknown samples for protein identification. The database allows text searches based on metadata such as experimental information of the samples, mass spectrometric instrument parameters or unique protein identifier like AGI codes. ProMEX integrates proteomics data with other levels of molecular organization including metabolite, pathway, and transcript information and may thus become a useful resource for plant systems biology studies. The ProMEX mass spectral library is available at http://promex.mpimp-golm.mpg.de/. Jan Hummel, Michaela Niemann, Stefanie Wienkoop, Waltraud X. Schulze, Dirk Steinhauser, Joachim Selbig, Dirk Walther 0001, Wolfram Weckwerth |
BMC Bioinform. | 6 |
| 2007 | Transcription factor target prediction using multiple short expression time series from Arabidopsis thalianaabstractBACKGROUND: The central role of transcription factors (TFs) in higher eukaryotes has led to much interest in deciphering transcriptional regulatory interactions. Even in the best case, experimental identification of TF target genes is error prone, and has been shown to be improved by considering additional forms of evidence such as expression data. Previous expression based methods have not explicitly tried to associate TFs with their targets and therefore largely ignored the treatment specific and time dependent nature of transcription regulation. RESULTS: In this study we introduce CERMT, Covariance based Extraction of Regulatory targets using Multiple Time series. Using simulated and real data we show that using multiple expression time series, selecting treatments in which the TF responds, allowing time shifts between TFs and their targets and using covariance to identify highly responding genes appear to be a good strategy. We applied our method to published TF - target gene relationships determined using expression profiling on TF mutants and show that in most cases we obtain significant target gene enrichment and in half of the cases this is sufficient to deliver a usable list of high-confidence target genes. CONCLUSION: CERMT could be immediately useful in refining possible target genes of candidate TFs using publicly available data, particularly for organisms lacking comprehensive TF binding data. In the future, we believe its incorporation with other forms of evidence may improve integrative genome-wide predictions of transcriptional networks. Henning Redestig, Daniel Weicht, Joachim Selbig, Matthew A. Hannah |
BMC Bioinform. | 3 |
| 2006 | Modelling Biological Networks by Action Languages Via Answer Set Programming
Susanne Grell, Torsten Schaub, Joachim Selbig |
ICLP | 3 |
| 2006 | Validation and functional annotation of expression-based clusters based on gene ontologyabstractBACKGROUND: The biological interpretation of large-scale gene expression data is one of the paramount challenges in current bioinformatics. In particular, placing the results in the context of other available functional genomics data, such as existing bio-ontologies, has already provided substantial improvement for detecting and categorizing genes of interest. One common approach is to look for functional annotations that are significantly enriched within a group or cluster of genes, as compared to a reference group. RESULTS: In this work, we suggest the information-theoretic concept of mutual information to investigate the relationship between groups of genes, as given by data-driven clustering, and their respective functional categories. Drawing upon related approaches (Gibbons and Roth, Genome Research 12:1574-1581, 2002), we seek to quantify to what extent individual attributes are sufficient to characterize a given group or cluster of genes. CONCLUSION: We show that the mutual information provides a systematic framework to assess the relationship between groups or clusters of genes and their functional annotations in a quantitative way. Within this framework, the mutual information allows us to address and incorporate several important issues, such as the interdependence of functional annotations and combinatorial combinations of attributes. It thus supplements and extends the conventional search for overrepresented attributes within a group or cluster of genes. In particular taking combinations of attributes into account, the mutual information opens the way to uncover specific functional descriptions of a group of genes or clustering result. All datasets and functional annotations used in this study are publicly available. All scripts used in the analysis are provided as additional files. Ralf Steuer, Peter Humburg, Joachim Selbig |
BMC Bioinform. | 3 |
| 2005 | Mtreemix: a software package for learning and using mixture models of mutagenetic treesabstractSUMMARY: Mixture models of mutagenetic trees constitute a class of probabilistic models for describing evolutionary processes that are characterized by the accumulation of permanent genetic changes. They have been applied to model the accumulation of chromosomal gains and losses in tumor development and the development of drug resistance-associated mutations in the HIV genome.Mtreemix is a software package for estimating mutagenetic trees mixture models from observed cross-sectional data and for using these models for predictions. We provide programs for model fitting, model selection, simulation, likelihood computation and waiting time estimation. AVAILABILITY: Mtreemix, including source code, documentation, sample data files and precompiled Solaris and Linux binaries, is freely available for non-commercial users at http://mtreemix.bioinf.mpi-sb.mpg.de/ Niko Beerenwinkel, Jörg Rahnenführer, Rolf Kaiser, Daniel Hoffmann, Joachim Selbig, Thomas Lengauer |
Bioinform. | 5 |
| 2005 | Computational methods for the design of effective therapies against drug resistant HIV strainsabstractThe development of drug resistance is a major obstacle to successful treatment of HIV infection. The extraordinary replication dynamics of HIV facilitates its escape from selective pressure exerted by the human immune system and by combination drug therapy. We have developed several computational methods whose combined use can support the design of optimal antiretroviral therapies based on viral genomic data. Niko Beerenwinkel, Tobias Sing, Thomas Lengauer, Jörg Rahnenführer, Kirsten Roomp, Igor Savenkov, Roman Fischer, Daniel Hoffmann, Joachim Selbig, Klaus Korn, Hauke Walter, Thomas Berg, Patrick Braun, Gerd Fätkenheuer, Mark Oette, Jürgen K. Rockstroh, Bernd Kupfer, Rolf Kaiser, Martin Däumer |
Bioinform. | 9 |
| 2005 | Non-linear PCA: a missing data approachabstractMOTIVATION: Visualizing and analysing the potential non-linear structure of a dataset is becoming an important task in molecular biology. This is even more challenging when the data have missing values. RESULTS: Here, we propose an inverse model that performs non-linear principal component analysis (NLPCA) from incomplete datasets. Missing values are ignored while optimizing the model, but can be estimated afterwards. Results are shown for both artificial and experimental datasets. In contrast to linear methods, non-linear methods were able to give better missing value estimations for non-linear structured data. APPLICATION: We applied this technique to a time course of metabolite data from a cold stress experiment on the model plant Arabidopsis thaliana, and could approximate the mapping function from any time point to the metabolite responses. Thus, the inverse NLPCA provides greatly improved information for better understanding the complex response to cold stress. CONTACT: [email protected]. Matthias Scholz, Fatma Kaplan, Charles L. Guy, Joachim Kopka, Joachim Selbig |
Bioinform. | 5 |
| 2005 | Species-specific analysis of protein sequence motifs using mutual informationabstractBACKGROUND: Protein sequence motifs are by definition short fragments of conserved amino acids, often associated with a specific function. Accordingly protein sequence profiles derived from multiple sequence alignments provide an alternative description of functional motifs characterizing families of related sequences. Such profiles conveniently reflect functional necessities by pointing out proximity at conserved sequence positions as well as depicting distances at variable positions. Discovering significant conservation characteristics within the variable positions of profiles mirrors group-specific and, in particular, evolutionary features of the underlying sequences. RESULTS: We describe the tool PROfile analysis based on Mutual Information (PROMI) that enables comparative analysis of user-classified protein sequences. PROMI is implemented as a web service using Perl and R as well as other publicly available packages and tools on the server-side. On the client-side platform-independence is achieved by generally applied internet delivery standards. As one possible application analysis of the zinc finger C2H2-type protein domain is introduced to illustrate the functionality of the tool. CONCLUSION: The web service PROMI should assist researchers to detect evolutionary correlations in protein profiles of defined biological sequences. It is available at http://promi.mpimp-golm.mpg.de where additional documentation can be found. Jan Hummel, Nima Keshvari, Wolfram Weckwerth, Joachim Selbig |
BMC Bioinform. | 4 |
| 2004 | Learning multiple evolutionary pathways from cross-sectional dataabstractWe introduce a mixture model of trees to describe evolutionary processes that are characterized by the accumulation of permanent genetic changes. The basic building block of the model is a directed weighted tree that generates a probability distribution on the set of all patterns of genetic events. We present an EM-like algorithm for learning a mixture model of K trees and show how to determine K with a maximum likelihood approach. As a case study we consider the accumulation of mutations in the HIV-1 reverse transcriptase that are associated with drug resistance. The fitted model is statistically validated as a density estimator and the stability of the model topology is analyzed. We obtain a generative probabilistic model for the development of drug resistance in HIV that agrees with biological knowledge. Further applications and extensions of the model are discussed. Niko Beerenwinkel, Jörg Rahnenführer, Martin Däumer, Daniel Hoffmann, Rolf Kaiser, Joachim Selbig, Thomas Lengauer |
RECOMB | 6 |
| 2004 | Threshold extraction in metabolite concentration dataabstractMOTIVATION: Continued development of analytical techniques based on gas chromatography and mass spectrometry now facilitates the generation of larger sets of metabolite concentration data. An important step towards the understanding of metabolite dynamics is the recognition of stable states where metabolite concentrations exhibit a simple behaviour. Such states can be characterized through the identification of significant thresholds in the concentrations. But general techniques for finding discretization thresholds in continuous data prove to be practically insufficient for detecting states due to the weak conditional dependences in concentration data. RESULTS: We introduce a method of recognizing states in the framework of decision tree induction. It is based upon a global analysis of decision forests where stability and quality are evaluated. It leads to the detection of thresholds that are both comprehensible and robust. Applied to metabolite concentration data, this method has led to the discovery of hidden states in the corresponding variables. Some of these reflect known properties of the biological experiments, and others point to putative new states. AVAILABILITY: An implementation of this approach can be obtained from the authors upon request. André Flöter, Jacques Nicolas, Torsten Schaub, Joachim Selbig |
Bioinform. | 4 |
| 2004 | PaVESy: Pathway Visualization and Editing SystemabstractUNLABELLED: A data managing system for editing and visualization of biological pathways is presented. The main component of PaVESy (Pathway Visualization and Editing System) is a relational SQL database system. The database design allows storage of biological objects, such as metabolites, proteins, genes and respective relations, which are required to assemble metabolic and regulatory biological interactions. The database model accommodates highly flexible annotation of biological objects by user-defined attributes. In addition, specific roles of objects are derived from these attributes in the context of user-defined interactions, e.g. in the course of pathway generation or during editing of the database content. Furthermore, the user may organize and arrange the database content within a folder structure and is free to group and annotate database objects of interest within customizable subsets. Thus, we allow an individualized view on the database content and facilitate user customization. A JAVA-based class library was developed, which serves as the database programming interface to PaVESy. This API provides classes, which implement the concepts of object persistence in SQL databases, such as entries, interactions, annotations, folders and subsets. We created editing and visualization tools for navigation in and visualization of the database content. User approved pathway assemblies are stored and may be retrieved for continued modification, annotation and export. Data export is interfaced with a range of network visualization programs, such as Pajek or other software allowing import of SBML or GML data format. AVAILABILITY: http://pavsey.mpimp-golm.mpg.de Alexander Lüdemann, Daniel Weicht, Joachim Selbig, Joachim Kopka |
Bioinform. | 3 |
| 2004 | Metabolite fingerprinting: detecting biological features by independent component analysisabstractMOTIVATION: Metabolite fingerprinting is a technology for providing information from spectra of total compositions of metabolites. Here, spectra acquisitions by microchip-based nanoflow-direct-infusion QTOF mass spectrometry, a simple and high throughput technique, is tested for its informative power. As a simple test case we are using Arabidopsis thaliana crosses. The question is how metabolite fingerprinting reflects the biological background. In many applications the classical principal component analysis (PCA) is used for detecting relevant information. Here a modern alternative is introduced-the independent component analysis (ICA). Due to its independence condition, ICA is more suitable for our questions than PCA. However, ICA has not been developed for a small number of high-dimensional samples, therefore a strategy is needed to overcome this limitation. RESULTS: To apply ICA successfully it is essential first to reduce the high dimension of the dataset, by using PCA. The number of principal components determines the quality of ICA significantly, therefore we propose a criterion for estimating the optimal dimension automatically. The kurtosis measure is used to order the extracted components to our interest. Applied to our A. thaliana data, ICA detects three relevant factors, two biological and one technical, and clearly outperforms the PCA. Matthias Scholz, S. Gatzek, Alisdair R. Fernie, Oliver Fiehn, Joachim Selbig |
Bioinform. | 5 |
| 2004 | Hypothesis-driven approach to predict transcriptional units from gene expression dataabstractMOTIVATION: A major issue in computational biology is the reconstruction of functional relationships among genes, for example the definition of regulatory or biochemical pathways. One step towards this aim is the elucidation of transcriptional units, which are characterized by co-responding changes in mRNA expression levels. These units of genes will allow the generation of hypotheses about respective functional interrelationships. Thus, the focus of analysis currently moves from well-established functional assignment through comparison of protein and DNA sequences towards analysis of transcriptional co-response. Tools that allow deducing common control of gene expression have the potential to complement and extend routine BLAST comparisons, because gene function may be inferred from common transcriptional control. RESULTS: We present a co-clustering strategy of genome sequence information and gene expression data, which was applied to identify transcriptional units within diverse compendia of expression profiles. The phenomenon of prokaryotic operons was selected as an ideal test case to generate well-founded hypotheses about transcriptional units. The existence of overlapping and ambiguous operon definitions allowed the investigation of constitutive and conditional expression of transcriptional units in independent gene expression experiments of Escherichia coli. Our approach allowed identification of operons with high accuracy. Furthermore, both constitutive mRNA co-response as well as conditional differences became apparent. Thus, we were able to generate insight into the possible biological relevance of gene co-response. We conclude that the suggested strategy will be amenable in general to the identification of transcriptional units beyond the chosen example of E.coli operons. AVAILABILITY: The analyses of E.coli transcript data presented here are available upon request or at http://csbdb.mpimp-golm.mpg.de/ Dirk Steinhauser, Björn H. Junker, Alexander Lüdemann, Joachim Selbig, Joachim Kopka |
Bioinform. | 4 |
| 2004 | Estimating mutual information using B-spline functions - an improved similarity measure for analysing gene expression dataabstractBACKGROUND: The information theoretic concept of mutual information provides a general framework to evaluate dependencies between variables. In the context of the clustering of genes with similar patterns of expression it has been suggested as a general quantity of similarity to extend commonly used linear measures. Since mutual information is defined in terms of discrete variables, its application to continuous data requires the use of binning procedures, which can lead to significant numerical errors for datasets of small or moderate size. RESULTS: In this work, we propose a method for the numerical estimation of mutual information from continuous data. We investigate the characteristic properties arising from the application of our algorithm and show that our approach outperforms commonly used algorithms: The significance, as a measure of the power of distinction from random correlation, is significantly increased. This concept is subsequently illustrated on two large-scale gene expression datasets and the results are compared to those obtained using other similarity measures.A C++ source code of our algorithm is available for non-commercial use from [email protected] upon request. CONCLUSION: The utilisation of mutual information as similarity measure enables the detection of non-linear correlations in gene expression datasets. Frequently applied linear correlation measures, which are often used on an ad-hoc basis without further justification, are thereby extended. Carsten O. Daub, Ralf Steuer, Joachim Selbig, Sebastian Kloska |
BMC Bioinform. | 3 |
| 2003 | MetaGeneAlyse: analysis of integrated transcriptional and metabolite dataabstractUNLABELLED: New techniques in sample preparation allow high throughput analysis of samples on the transcriptional as well as on the metabolic level. We present a service accessible via the web that allows the analysis of integrated data sets that combine gene-expression data and metabolic data. After uploading, data sets can be normalized, clustered by various methods and results can be graphically visualized. All calculations are carried out on a server, so even time- and memory-consuming analyses can be done independently of the performance of the client. AVAILABILITY: The service is accessible via web-interface at http://metagenealyse.mpimp-golm.mpg.de/ Carsten O. Daub, Sebastian Kloska, Joachim Selbig |
Bioinform. | 3 |
| 1999 | Decision tree-based formation of consensus protein secondary structure predictionabstractAbstract Motivation: Prediction of protein secondary structure provides information that is useful for other prediction methods like fold recognition and ab initio 3D prediction. A consensus prediction constructed from the output of several methods should yield more reliable results than each of the individual methods. Method: We present an approach that reveals subtle but systematic differences in the output of different secondary structure prediction methods allowing the derivation of coherent consensus predictions. The method uses a machine learning technique that builds decision trees from existing data. Results: The first results of our analysis show that consensus prediction of protein secondary structure may be improved both quantitatively and qualitatively. Availability: Our method for consensus secondary structure prediction CoDe (Consensus formation by Decision tree learning) is based on machine learning and will be integrated into the ToPLign system (Toolbox for Protein aLignment) which can be accessed at http://cartan.gmd.de/ToPLign.html. Contact: {mevissen,selbig}@gmd.de Joachim Selbig, Heinz-Theodor Mevissen, Thomas Lengauer |
Bioinform. | 1 |
| 1981 | Concept Learning by Structured Examples - An Algebraic Approach
Fritz Wysotzki, Werner Kolbe, Joachim Selbig |
IJCAI | 3 |