VLDB 2026 Research / reviewers in the wild / expert
Jens Nielsen
dblp:03/6254
· DBLP profile ↗
20ranked-venue papers
0as first author
4since 2021 · last 2023
0000-0002-9955-6003ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 19 · 4 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | HGTphyloDetect: facilitating the identification and phylogenetic analysis of horizontal gene transferabstractBACKGROUND: Horizontal gene transfer (HGT) is an important driver in genome evolution, gain-of-function, and metabolic adaptation to environmental niches. Genome-wide identification of putative HGT events has become increasingly practical, given the rapid growth of genomic data. However, existing HGT analysis toolboxes are not widely used, limited by their inability to perform phylogenetic reconstruction to explore potential donors, and the detection of HGT from both evolutionarily distant and closely related species. RESULTS: In this study, we have developed HGTphyloDetect, which is a versatile computational toolbox that combines high-throughput analysis with phylogenetic inference, to facilitate comprehensive investigation of HGT events. Two case studies with Saccharomyces cerevisiae and Candida versatilis demonstrate the ability of HGTphyloDetect to identify horizontally acquired genes with high accuracy. In addition, HGTphyloDetect enables phylogenetic analysis to illustrate a likely path of gene transmission among the evolutionarily distant or closely related species. CONCLUSIONS: The HGTphyloDetect computational toolbox is designed for ease of use and can accurately find HGT events with a very low false discovery rate in a high-throughput manner. The HGTphyloDetect toolbox and its related user tutorial are freely available at https://github.com/SysBioChalmers/HGTphyloDetect. Le Yuan, Hongzhong Lu, Feiran Li, Jens Nielsen, Eduard J. Kerkhoven |
Briefings Bioinform. | 4 |
| 2021 | Addressing the heterogeneity in liver diseases using biological networksabstractThe abnormalities in human metabolism have been implicated in the progression of several complex human diseases, including certain cancers. Hence, deciphering the underlying molecular mechanisms associated with metabolic reprogramming in a disease state can greatly assist in elucidating the disease aetiology. An invaluable tool for establishing connections between global metabolic reprogramming and disease development is the genome-scale metabolic model (GEM). Here, we review recent work on the reconstruction of cell/tissue-type and cancer-specific GEMs and their use in identifying metabolic changes occurring in response to liver disease development, stratification of the heterogeneous disease population and discovery of novel drug targets and biomarkers. We also discuss how GEMs can be integrated with other biological networks for generating more comprehensive cell/tissue models. In addition, we review the various biological network analyses that have been employed for the development of efficient treatment strategies. Finally, we present three case studies in which independent studies converged on conclusions underlying liver disease. Simon Lam, Stephen Doran, Hatice Hilal Yuksel, Ozlem Altay, Hasan Turkez, Jens Nielsen, Jan Boren, Mathias Uhlen, Adil Mardinoglu |
Briefings Bioinform. | 6 |
| 2021 | A novel yeast hybrid modeling framework integrating Boolean and enzyme-constrained networks enables exploration of the interplay between signaling and metabolismabstractThe interplay between nutrient-induced signaling and metabolism plays an important role in maintaining homeostasis and its malfunction has been implicated in many different human diseases such as obesity, type 2 diabetes, cancer, and neurological disorders. Therefore, unraveling the role of nutrients as signaling molecules and metabolites together with their interconnectivity may provide a deeper understanding of how these conditions occur. Both signaling and metabolism have been extensively studied using various systems biology approaches. However, they are mainly studied individually and in addition, current models lack both the complexity of the dynamics and the effects of the crosstalk in the signaling system. To gain a better understanding of the interconnectivity between nutrient signaling and metabolism in yeast cells, we developed a hybrid model, combining a Boolean module, describing the main pathways of glucose and nitrogen signaling, and an enzyme-constrained model accounting for the central carbon metabolism of Saccharomyces cerevisiae, using a regulatory network as a link. The resulting hybrid model was able to capture a diverse utalization of isoenzymes and to our knowledge outperforms constraint-based models in the prediction of individual enzymes for both respiratory and mixed metabolism. The model showed that during fermentation, enzyme utilization has a major contribution in governing protein allocation, while in low glucose conditions robustness and control are prioritized. In addition, the model was capable of reproducing the regulatory effects that are associated with the Crabtree effect and glucose repression, as well as regulatory effects associated with lifespan increase during caloric restriction. Overall, we show that our hybrid model provides a comprehensive framework for the study of the non-trivial effects of the interplay between signaling and metabolism, suggesting connections between the Snf1 signaling pathways and processes that have been related to chronological lifespan of yeast cells. Linnea Österberg, Iván Domenzain, Julia Münch, Jens Nielsen, Stefan Hohmann, Marija Cvijovic |
PLoS Comput. Biol. | 4 |
| 2021 | Machine learning-based investigation of the cancer protein secretory pathwayabstractDeregulation of the protein secretory pathway (PSP) is linked to many hallmarks of cancer, such as promoting tissue invasion and modulating cell-cell signaling. The collection of secreted proteins processed by the PSP, known as the secretome, is often studied due to its potential as a reservoir of tumor biomarkers. However, there has been less focus on the protein components of the secretory machinery itself. We therefore investigated the expression changes in secretory pathway components across many different cancer types. Specifically, we implemented a dual approach involving differential expression analysis and machine learning to identify PSP genes whose expression was associated with key tumor characteristics: mutation of p53, cancer status, and tumor stage. Eight different machine learning algorithms were included in the analysis to enable comparison between methods and to focus on signals that were robust to algorithm type. The machine learning approach was validated by identifying PSP genes known to be regulated by p53, and even outperformed the differential expression analysis approach. Among the different analysis methods and cancer types, the kinesin family members KIF20A and KIF23 were consistently among the top genes associated with malignant transformation or tumor stage. However, unlike most cancer types which exhibited elevated KIF20A expression that remained relatively constant across tumor stages, renal carcinomas displayed a more gradual increase that continued with increasing disease severity. Collectively, our study demonstrates the complementary nature of a combined differential expression and machine learning approach for analyzing gene expression data, and highlights key PSP components relevant to features of tumor pathophysiology that may constitute potential therapeutic targets. Rasool Saghaleyni, Azam Sheikh Muhammad, Pramod Bangalore, Jens Nielsen, Jonathan L. Robinson |
PLoS Comput. Biol. | 4 |
| 2019 | Reconstruction and analysis of a Kluyveromyces marxianus genome-scale metabolic modelabstractBACKGROUND: Kluyveromyces marxianus is a thermotolerant yeast with multiple biotechnological potentials for industrial applications, which can metabolize a broad range of carbon sources, including less conventional sugars like lactose, xylose, arabinose and inulin. These phenotypic traits are sustained even up to 45 °C, what makes it a relevant candidate for industrial biotechnology applications, such as ethanol production. It is therefore of much interest to get more insight into the metabolism of this yeast. Recent studies suggested, that thermotolerance is achieved by reducing the number of growth-determining proteins or suppressing oxidative phosphorylation. Here we aimed to find related factors contributing to the thermotolerance of K. marxianus. RESULTS: Here, we reported the first genome-scale metabolic model of Kluyveromyces marxianus, iSM996, using a publicly available Kluyveromyces lactis model as template. The model was manually curated and refined to include the missing species-specific metabolic capabilities. The iSM996 model includes 1913 reactions, associated with 996 genes and 1531 metabolites. It performed well to predict the carbon source utilization and growth rates under different growth conditions. Moreover, the model was coupled with transcriptomics data and used to perform simulations at various growth temperatures. CONCLUSIONS: K. marxianus iSM996 represents a well-annotated metabolic model of thermotolerant yeast, which provides a new insight into theoretical metabolic profiles at different temperatures of K. marxianus. This could accelerate the integrative analysis of multi-omics data, leading to model-driven strain design and improvement. Simonas Marcisauskas, Boyang Ji, Jens Nielsen |
BMC Bioinform. | 3 |
| 2018 | RAVEN 2.0: A versatile toolbox for metabolic network reconstruction and a case study on Streptomyces coelicolorabstractRAVEN is a commonly used MATLAB toolbox for genome-scale metabolic model (GEM) reconstruction, curation and constraint-based modelling and simulation. Here we present RAVEN Toolbox 2.0 with major enhancements, including: (i) de novo reconstruction of GEMs based on the MetaCyc pathway database; (ii) a redesigned KEGG-based reconstruction pipeline; (iii) convergence of reconstructions from various sources; (iv) improved performance, usability, and compatibility with the COBRA Toolbox. Capabilities of RAVEN 2.0 are here illustrated through de novo reconstruction of GEMs for the antibiotic-producing bacterium Streptomyces coelicolor. Comparison of the automated de novo reconstructions with the iMK1208 model, a previously published high-quality S. coelicolor GEM, exemplifies that RAVEN 2.0 can capture most of the manually curated model. The generated de novo reconstruction is subsequently used to curate iMK1208 resulting in Sco4, the most comprehensive GEM of S. coelicolor, with increased coverage of both primary and secondary metabolism. This increased coverage allows the use of Sco4 to predict novel genome editing targets for optimized secondary metabolites production. As such, we demonstrate that RAVEN 2.0 can be used not only for de novo GEM reconstruction, but also for curating existing models based on up-to-date databases. Both RAVEN 2.0 and Sco4 are distributed through GitHub to facilitate usage and further development by the community (https://github.com/SysBioChalmers/RAVEN and https://github.com/SysBioChalmers/Streptomyces_coelicolor-GEM). Hao Wang 0078, Simonas Marcisauskas, Benjamín J. Sánchez, Iván Domenzain, Daniel Hermansson, Rasmus Agren, Jens Nielsen, Eduard J. Kerkhoven |
PLoS Comput. Biol. | 7 |
| 2017 | Systematic inference of functional phosphorylation events in yeast metabolismabstractMOTIVATION: Protein phosphorylation is a post-translational modification that affects proteins by changing their structure and conformation in a rapid and reversible way, and it is an important mechanism for metabolic regulation in cells. Phosphoproteomics enables high-throughput identification of phosphorylation events on metabolic enzymes, but identifying functional phosphorylation events still requires more detailed biochemical characterization. Therefore, development of computational methods for investigating unknown functions of a large number of phosphorylation events identified by phosphoproteomics has received increased attention. RESULTS: We developed a mathematical framework that describes the relationship between phosphorylation level of a metabolic enzyme and the corresponding flux through the enzyme. Using this framework, it is possible to quantitatively estimate contribution of phosphorylation events to flux changes. We showed that phosphorylation regulation analysis, combined with a systematic workflow and correlation analysis, can be used for inference of functional phosphorylation events in steady and dynamic conditions, respectively. Using this analysis, we assigned functionality to phosphorylation events of 17 metabolic enzymes in the yeast Saccharomyces cerevisiae , among which 10 are novel. Phosphorylation regulation analysis cannot only be extended for inference of other functional post-translational modifications but also be a promising scaffold for multi-omics data integration in systems biology. AVAILABILITY AND IMPLEMENTATION: Matlab codes for flux balance analysis in this study are available in Supplementary material. CONTACT: [email protected] or [email protected]. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yu Chen 0095, Jens Nielsen |
Bioinform. | 3 |
| 2015 | Logical transformation of genome-scale metabolic models for gene level applications and analysisabstractMOTIVATION: In recent years, genome-scale metabolic models (GEMs) have played important roles in areas like systems biology and bioinformatics. However, because of the complexity of gene-reaction associations, GEMs often have limitations in gene level analysis and related applications. Hence, the existing methods were mainly focused on applications and analysis of reactions and metabolites. RESULTS: Here, we propose a framework named logic transformation of model (LTM) that is able to simplify the gene-reaction associations and enables integration with other developed methods for gene level applications. We show that the transformed GEMs have increased reaction and metabolite number as well as degree of freedom in flux balance analysis, but the gene-reaction associations and the main features of flux distributions remain constant. In addition, we develop two methods, OptGeneKnock and FastGeneSL by combining LTM with previously developed reaction-based methods. We show that the FastGeneSL outperforms exhaustive search. Finally, we demonstrate the use of the developed methods in two different case studies. We could design fast genetic intervention strategies for targeted overproduction of biochemicals and identify double and triple synthetic lethal gene sets for inhibition of hepatocellular carcinoma tumor growth through the use of OptGeneKnock and FastGeneSL, respectively. AVAILABILITY AND IMPLEMENTATION: Source code implemented in MATLAB, RAVEN toolbox and COBRA toolbox, is public available at https://sourceforge.net/projects/logictransformationofmodel. Cheng Zhang 0002, Boyang Ji, Adil Mardinoglu, Jens Nielsen, Qiang Hua |
Bioinform. | 4 |
| 2015 | Metabolic Needs and Capabilities of Toxoplasma gondii through Combined Computational and Experimental AnalysisabstractToxoplasma gondii is a human pathogen prevalent worldwide that poses a challenging and unmet need for novel treatment of toxoplasmosis. Using a semi-automated reconstruction algorithm, we reconstructed a genome-scale metabolic model, ToxoNet1. The reconstruction process and flux-balance analysis of the model offer a systematic overview of the metabolic capabilities of this parasite. Using ToxoNet1 we have identified significant gaps in the current knowledge of Toxoplasma metabolic pathways and have clarified its minimal nutritional requirements for replication. By probing the model via metabolic tasks, we have further defined sets of alternative precursors necessary for parasite growth. Within a human host cell environment, ToxoNet1 predicts a minimal set of 53 enzyme-coding genes and 76 reactions to be essential for parasite replication. Double-gene-essentiality analysis identified 20 pairs of genes for which simultaneous deletion is deleterious. To validate several predictions of ToxoNet1 we have performed experimental analyses of cytosolic acetyl-CoA biosynthesis. ATP-citrate lyase and acetyl-CoA synthase were localised and their corresponding genes disrupted, establishing that each of these enzymes is dispensable for the growth of T. gondii, however together they make a synthetic lethal pair. Stepan Tymoshenko, Rebecca D. Oppenheim, Rasmus Agren, Jens Nielsen, Dominique Soldati-Favre, Vassily Hatzimanikatis |
PLoS Comput. Biol. | 4 |
| 2014 | Kiwi: a tool for integration and visualization of network topology and gene-set analysisabstractBACKGROUND: The analysis of high-throughput data in biology is aided by integrative approaches such as gene-set analysis. Gene-sets can represent well-defined biological entities (e.g. metabolites) that interact in networks (e.g. metabolic networks), to exert their function within the cell. Data interpretation can benefit from incorporating the underlying network, but there are currently no optimal methods that link gene-set analysis and network structures. RESULTS: Here we present Kiwi, a new tool that processes output data from gene-set analysis and integrates them with a network structure such that the inherent connectivity between gene-sets, i.e. not simply the gene overlap, becomes apparent. In two case studies, we demonstrate that standard gene-set analysis points at metabolites regulated in the interrogated condition. Nevertheless, only the integration of the interactions between these metabolites provides an extra layer of information that highlights how they are tightly connected in the metabolic network. CONCLUSIONS: Kiwi is a tool that enhances interpretability of high-throughput data. It allows the users not only to discover a list of significant entities or processes as in gene-set analysis, but also to visualize whether these entities or processes are isolated or connected by means of their biological interaction. Kiwi is available as a Python package at http://www.sysbio.se/kiwi and an online tool in the BioMet Toolbox at http://www.biomet-toolbox.org. Leif Väremo, Francesco Gatto 0001, Jens Nielsen |
BMC Bioinform. | 3 |
| 2014 | Metagenomic Data Utilization and Analysis (MEDUSA) and Construction of a Global Gut Microbial Gene CatalogueabstractMetagenomic sequencing has contributed important new knowledge about the microbes that live in a symbiotic relationship with humans. With modern sequencing technology it is possible to generate large numbers of sequencing reads from a metagenome but analysis of the data is challenging. Here we present the bioinformatics pipeline MEDUSA that facilitates analysis of metagenomic reads at the gene and taxonomic level. We also constructed a global human gut microbial gene catalogue by combining data from 4 studies spanning 3 continents. Using MEDUSA we mapped 782 gut metagenomes to the global gene catalogue and a catalogue of sequenced microbial species. Hereby we find that all studies share about half a million genes and that on average 300,000 genes are shared by half the studied subjects. The gene richness is higher in the European studies compared to Chinese and American and this is also reflected in the species richness. Even though it is possible to identify common species and a core set of genes, we find that there are large variations in abundance of species and genes. Fredrik H. Karlsson, Intawat Nookaew, Jens Nielsen |
PLoS Comput. Biol. | 3 |
| 2013 | FANTOM: Functional and taxonomic analysis of metagenomesabstractBACKGROUND: Interpretation of quantitative metagenomics data is important for our understanding of ecosystem functioning and assessing differences between various environmental samples. There is a need for an easy to use tool to explore the often complex metagenomics data in taxonomic and functional context. RESULTS: Here we introduce FANTOM, a tool that allows for exploratory and comparative analysis of metagenomics abundance data integrated with metadata information and biological databases. Importantly, FANTOM can make use of any hierarchical database and it comes supplied with NCBI taxonomic hierarchies as well as KEGG Orthology, COG, PFAM and TIGRFAM databases. CONCLUSIONS: The software is implemented in Python, is platform independent, and is available at http://www.sysbio.se/Fantom. Kemal Sanli, Fredrik H. Karlsson, Intawat Nookaew, Jens Nielsen |
BMC Bioinform. | 4 |
| 2013 | The RAVEN Toolbox and Its Use for Generating a Genome-scale Metabolic Model for Penicillium chrysogenumabstractWe present the RAVEN (Reconstruction, Analysis and Visualization of Metabolic Networks) Toolbox: a software suite that allows for semi-automated reconstruction of genome-scale models. It makes use of published models and/or the KEGG database, coupled with extensive gap-filling and quality control features. The software suite also contains methods for visualizing simulation results and omics data, as well as a range of methods for performing simulations and analyzing the results. The software is a useful tool for system-wide data analysis in a metabolic context and for streamlined reconstruction of metabolic networks based on protein homology. The RAVEN Toolbox workflow was applied in order to reconstruct a genome-scale metabolic model for the important microbial cell factory Penicillium chrysogenum Wisconsin54-1255. The model was validated in a bibliomic study of in total 440 references, and it comprises 1471 unique biochemical reactions and 1006 ORFs. It was then used to study the roles of ATP and NADPH in the biosynthesis of penicillin, and to identify potential metabolic engineering targets for maximization of penicillin production. Rasmus Agren, Saeed Shoaie, Wanwipa Vongsangnak, Intawat Nookaew, Jens Nielsen |
PLoS Comput. Biol. | 6 |
| 2013 | Quantitative Analysis of Glycerol Accumulation, Glycolysis and Growth under Hyper Osmotic StressabstractWe provide an integrated dynamic view on a eukaryotic osmolyte system, linking signaling with regulation of gene expression, metabolic control and growth. Adaptation to osmotic changes enables cells to adjust cellular activity and turgor pressure to an altered environment. The yeast Saccharomyces cerevisiae adapts to hyperosmotic stress by activating the HOG signaling cascade, which controls glycerol accumulation. The Hog1 kinase stimulates transcription of genes encoding enzymes required for glycerol production (Gpd1, Gpp2) and glycerol import (Stl1) and activates a regulatory enzyme in glycolysis (Pfk26/27). In addition, glycerol outflow is prevented by closure of the Fps1 glycerol facilitator. In order to better understand the contributions to glycerol accumulation of these different mechanisms and how redox and energy metabolism as well as biomass production are maintained under such conditions we collected an extensive dataset. Over a period of 180 min after hyperosmotic shock we monitored in wild type and different mutant cells the concentrations of key metabolites and proteins relevant for osmoadaptation. The dataset was used to parameterize an ODE model that reproduces the generated data very well. A detailed computational analysis using time-dependent response coefficients showed that Pfk26/27 contributes to rerouting glycolytic flux towards lower glycolysis. The transient growth arrest following hyperosmotic shock further adds to redirecting almost all glycolytic flux from biomass towards glycerol production. Osmoadaptation is robust to loss of individual adaptation pathways because of the existence and upregulation of alternative routes of glycerol accumulation. For instance, the Stl1 glycerol importer contributes to glycerol accumulation in a mutant with diminished glycerol production capacity. In addition, our observations suggest a role for trehalose accumulation in osmoadaptation and that Hog1 probably directly contributes to the regulation of the Fps1 glycerol facilitator. Taken together, we elucidated how different metabolic adaptation mechanisms cooperate and provide hypotheses for further experimental studies. Elzbieta Petelenz-Kurdziel, Clemens Kühn, Bodil Nordlander, Dagmara Klein, Kuk-Ki Hong, Therese Jacobson, Peter Dahl, Jörg Schaber, Jens Nielsen, Stefan Hohmann, Edda Klipp |
PLoS Comput. Biol. | 9 |
| 2012 | Reconstruction of Genome-Scale Active Metabolic Networks for 69 Human Cell Types and 16 Cancer Types Using INITabstractDevelopment of high throughput analytical methods has given physicians the potential access to extensive and patient-specific data sets, such as gene sequences, gene expression profiles or metabolite footprints. This opens for a new approach in health care, which is both personalized and based on system-level analysis. Genome-scale metabolic networks provide a mechanistic description of the relationships between different genes, which is valuable for the analysis and interpretation of large experimental data-sets. Here we describe the generation of genome-scale active metabolic networks for 69 different cell types and 16 cancer types using the INIT (Integrative Network Inference for Tissues) algorithm. The INIT algorithm uses cell type specific information about protein abundances contained in the Human Proteome Atlas as the main source of evidence. The generated models constitute the first step towards establishing a Human Metabolic Atlas, which will be a comprehensive description (accessible online) of the metabolism of different human cell types, and will allow for tissue-level and organism-level simulations in order to achieve a better understanding of complex diseases. A comparative analysis between the active metabolic networks of cancer types and healthy cell types allowed for identification of cancer-specific metabolic features that constitute generic potential drug targets for cancer treatment. Rasmus Agren, Sergio Bordel, Adil Mardinoglu, Natapol Pornputtapong, Intawat Nookaew, Jens Nielsen |
PLoS Comput. Biol. | 6 |
| 2010 | Sampling the Solution Space in Genome-Scale Metabolic Networks Reveals Transcriptional Regulation in Key EnzymesabstractGenome-scale metabolic models are available for an increasing number of organisms and can be used to define the region of feasible metabolic flux distributions. In this work we use as constraints a small set of experimental metabolic fluxes, which reduces the region of feasible metabolic states. Once the region of feasible flux distributions has been defined, a set of possible flux distributions is obtained by random sampling and the averages and standard deviations for each of the metabolic fluxes in the genome-scale model are calculated. These values allow estimation of the significance of change for each reaction rate between different conditions and comparison of it with the significance of change in gene transcription for the corresponding enzymes. The comparison of flux change and gene expression allows identification of enzymes showing a significant correlation between flux change and expression change (transcriptional regulation) as well as reactions whose flux change is likely to be driven only by changes in the metabolite concentrations (metabolic regulation). The changes due to growth on four different carbon sources and as a consequence of five gene deletions were analyzed for Saccharomyces cerevisiae. The enzymes with transcriptional regulation showed enrichment in certain transcription factors. This has not been previously reported. The information provided by the presented method could guide the discovery of new metabolic engineering strategies or the identification of drug targets for treatment of metabolic diseases. Sergio Bordel, Rasmus Agren, Jens Nielsen |
PLoS Comput. Biol. | 3 |
| 2008 | Natural computation meta-heuristics for the in silico optimization of microbial strainsabstractBACKGROUND: One of the greatest challenges in Metabolic Engineering is to develop quantitative models and algorithms to identify a set of genetic manipulations that will result in a microbial strain with a desirable metabolic phenotype which typically means having a high yield/productivity. This challenge is not only due to the inherent complexity of the metabolic and regulatory networks, but also to the lack of appropriate modelling and optimization tools. To this end, Evolutionary Algorithms (EAs) have been proposed for in silico metabolic engineering, for example, to identify sets of gene deletions towards maximization of a desired physiological objective function. In this approach, each mutant strain is evaluated by resorting to the simulation of its phenotype using the Flux-Balance Analysis (FBA) approach, together with the premise that microorganisms have maximized their growth along natural evolution. RESULTS: This work reports on improved EAs, as well as novel Simulated Annealing (SA) algorithms to address the task of in silico metabolic engineering. Both approaches use a variable size set-based representation, thereby allowing the automatic finding of the best number of gene deletions necessary for achieving a given productivity goal. The work presents extensive computational experiments, involving four case studies that consider the production of succinic and lactic acid as the targets, by using S. cerevisiae and E. coli as model organisms. The proposed algorithms are able to reach optimal/near-optimal solutions regarding the production of the desired compounds and presenting low variability among the several runs. CONCLUSION: The results show that the proposed SA and EA both perform well in the optimization task. A comparison between them is favourable to the SA in terms of consistency in obtaining optimal solutions and faster convergence. In both cases, the use of variable size representations allows the automatic discovery of the approximate number of gene deletions, without compromising the optimality of the solutions. Miguel Rocha 0001, Paulo Maia, Rui Mendes 0001, José P. Pinto, Eugénio C. Ferreira, Jens Nielsen, Kiran Raosaheb Patil, Isabel Rocha |
BMC Bioinform. | 6 |
| 2006 | Robust multi-scale clustering of large DNA microarray datasets with the consensus algorithmabstractMOTIVATION: Hierarchical and relocation clustering (e.g. K-means and self-organizing maps) have been successful tools in the display and analysis of whole genome DNA microarray expression data. However, the results of hierarchical clustering are sensitive to outliers, and most relocation methods give results which are dependent on the initialization of the algorithm. Therefore, it is difficult to assess the significance of the results. We have developed a consensus clustering algorithm, where the final result is averaged over multiple clustering runs, giving a robust and reproducible clustering, capable of capturing small signal variations. The algorithm preserves valuable properties of hierarchical clustering, which is useful for visualization and interpretation of the results. RESULTS: We show for the first time that one can take advantage of multiple clustering runs in DNA microarray analysis by collecting re-occurring clustering patterns in a co-occurrence matrix. The results show that consensus clustering obtained from clustering multiple times with Variational Bayes Mixtures of Gaussians or K-means significantly reduces the classification error rate for a simulated dataset. The method is flexible and it is possible to find consensus clusters from different clustering algorithms. Thus, the algorithm can be used as a framework to test in a quantitative manner the homogeneity of different clustering algorithms. We compare the method with a number of state-of-the-art clustering methods. It is shown that the method is robust and gives low classification error rates for a realistic, simulated dataset. The algorithm is also demonstrated for real datasets. It is shown that more biological meaningful transcriptional patterns can be found without conservative statistical or fold-change exclusion of data. AVAILABILITY: Matlab source code for the clustering algorithm ClusterLustre, and the simulated dataset for testing are available upon request from T.G. and O.W. Thomas Grotkjær, Ole Winther, Birgitte Regenberg, Jens Nielsen, Lars Kai Hansen |
Bioinform. | 4 |
| 2005 | Evolutionary programming as a platform for in silico metabolic engineeringabstractBACKGROUND: Through genetic engineering it is possible to introduce targeted genetic changes and hereby engineer the metabolism of microbial cells with the objective to obtain desirable phenotypes. However, owing to the complexity of metabolic networks, both in terms of structure and regulation, it is often difficult to predict the effects of genetic modifications on the resulting phenotype. Recently genome-scale metabolic models have been compiled for several different microorganisms where structural and stoichiometric complexity is inherently accounted for. New algorithms are being developed by using genome-scale metabolic models that enable identification of gene knockout strategies for obtaining improved phenotypes. However, the problem of finding optimal gene deletion strategy is combinatorial and consequently the computational time increases exponentially with the size of the problem, and it is therefore interesting to develop new faster algorithms. RESULTS: In this study we report an evolutionary programming based method to rapidly identify gene deletion strategies for optimization of a desired phenotypic objective function. We illustrate the proposed method for two important design parameters in industrial fermentations, one linear and other non-linear, by using a genome-scale model of the yeast Saccharomyces cerevisiae. Potential metabolic engineering targets for improved production of succinic acid, glycerol and vanillin are identified and underlying flux changes for the predicted mutants are discussed. CONCLUSION: We show that evolutionary programming enables solving large gene knockout problems in relatively short computational time. The proposed algorithm also allows the optimization of non-linear objective functions or incorporation of non-linear constraints and additionally provides a family of close to optimal solutions. The identified metabolic engineering strategies suggest that non-intuitive genetic modifications span several different pathways and may be necessary for solving challenging metabolic engineering problems. Kiran Raosaheb Patil, Isabel Rocha, Jochen Förster, Jens Nielsen |
BMC Bioinform. | 4 |
| 2002 | Short trees in polygons
Pawel Winter, Martin Zachariasen, Jens Nielsen |
Discret. Appl. Math. | 3 |