VLDB 2026 Research / reviewers in the wild / expert
Eivind Almaas
dblp:87/6946
· DBLP profile ↗
13ranked-venue papers
1as first author
5since 2021 · last 2023
0000-0002-9125-326XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Genome-scale metabolic models reveal determinants of phenotypic differences in non-Saccharomyces yeastsabstractBACKGROUND: Use of alternative non-Saccharomyces yeasts in wine and beer brewing has gained more attention the recent years. This is both due to the desire to obtain a wider variety of flavours in the product and to reduce the final alcohol content. Given the metabolic differences between the yeast species, we wanted to account for some of the differences by using in silico models. RESULTS: We created and studied genome-scale metabolic models of five different non-Saccharomyces species using an automated processes. These were: Metschnikowia pulcherrima, Lachancea thermotolerans, Hanseniaspora osmophila, Torulaspora delbrueckii and Kluyveromyces lactis. Using the models, we predicted that M. pulcherrima, when compared to the other species, conducts more respiration and thus produces less fermentation products, a finding which agrees with experimental data. Complex I of the electron transport chain was to be present in M. pulcherrima, but absent in the others. The predicted importance of Complex I was diminished when we incorporated constraints on the amount of enzymatic protein, as this shifts the metabolism towards fermentation. CONCLUSIONS: Our results suggest that Complex I in the electron transport chain is a key differentiator between Metschnikowia pulcherrima and the other yeasts considered. Yet, more annotations and experimental data have the potential to improve model quality in order to increase fidelity and confidence in these results. Further experiments should be conducted to confirm the in vivo effect of Complex I in M. pulcherrima and its respiratory metabolism. Jakob P. Pettersen, Sandra Castillo, Paula Jouhten, Eivind Almaas |
BMC Bioinform. | 4 |
| 2022 | csdR, an R package for differential co-expression analysisabstractBACKGROUND: Differential co-expression network analysis has become an important tool to gain understanding of biological phenotypes and diseases. The CSD algorithm is a method to generate differential co-expression networks by comparing gene co-expressions from two different conditions. Each of the gene pairs is assigned conserved (C), specific (S) and differentiated (D) scores based on the co-expression of the gene pair between the two conditions. The result of the procedure is a network where the nodes are genes and the links are the gene pairs with the highest C-, S-, and D-scores. However, the existing CSD-implementations suffer from poor computational performance, difficult user procedures and lack of documentation. RESULTS: We created the R-package csdR aimed at reaching good performance together with ease of use, sufficient documentation, and with the ability to play well with other tools for data analysis. csdR was benchmarked on a realistic dataset with 20,645 genes. After verifying that the chosen number of iterations gave sufficient robustness, we tested the performance against the two existing CSD implementations. csdR was superior in performance to one of the implementations, whereas the other did not run. Our implementation can utilize multiple processing cores. However, we were unable to achieve more than [Formula: see text]2.7 parallel speedup with saturation reached at about 10 cores. CONCLUSION: The results suggest that csdR is a useful tool for differential co-expression analysis and is able to generate robust results within a workday on datasets of realistic sizes when run on a workstation or compute server. Jakob P. Pettersen, Eivind Almaas |
BMC Bioinform. | 2 |
| 2021 | Automatic reconstruction of metabolic pathways from identified biosynthetic gene clustersabstractBACKGROUND: A wide range of bioactive compounds is produced by enzymes and enzymatic complexes encoded in biosynthetic gene clusters (BGCs). These BGCs can be identified and functionally annotated based on their DNA sequence. Candidates for further research and development may be prioritized based on properties such as their functional annotation, (dis)similarity to known BGCs, and bioactivity assays. Production of the target compound in the native strain is often not achievable, rendering heterologous expression in an optimized host strain as a promising alternative. Genome-scale metabolic models are frequently used to guide strain development, but large-scale incorporation and testing of heterologous production of complex natural products in this framework is hampered by the amount of manual work required to translate annotated BGCs to metabolic pathways. To this end, we have developed a pipeline for an automated reconstruction of BGC associated metabolic pathways responsible for the synthesis of non-ribosomal peptides and polyketides, two of the dominant classes of bioactive compounds. RESULTS: The developed pipeline correctly predicts 72.8% of the metabolic reactions in a detailed evaluation of 8 different BGCs comprising 228 functional domains. By introducing the reconstructed pathways into a genome-scale metabolic model we demonstrate that this level of accuracy is sufficient to make reliable in silico predictions with respect to production rate and gene knockout targets. Furthermore, we apply the pipeline to a large BGC database and reconstruct 943 metabolic pathways. We identify 17 enzymatic reactions using high-throughput assessment of potential knockout targets for increasing the production of any of the associated compounds. However, the targets only provide a relative increase of up to 6% compared to wild-type production rates. CONCLUSION: With this pipeline we pave the way for an extended use of genome-scale metabolic models in strain design of heterologous expression hosts. In this context, we identified generic knockout targets for the increased production of heterologous compounds. However, as the predicted increase is minor for any of the single-reaction knockout targets, these results indicate that more sophisticated strain-engineering strategies are necessary for the development of efficient BGC expression hosts. Snorre Sulheim, Fredrik A. Fossheim, Alexander Wentzel, Eivind Almaas |
BMC Bioinform. | 4 |
| 2021 | Correction to: Automatic reconstruction of metabolic pathways from identified biosynthetic gene clustersabstractAn amendment to this paper has been published and can be accessed via the original article. Snorre Sulheim, Fredrik A. Fossheim, Alexander Wentzel, Eivind Almaas |
BMC Bioinform. | 4 |
| 2021 | Genome-scale metabolic modelling when changes in environmental conditions affect biomass compositionabstractGenome-scale metabolic modeling is an important tool in the study of metabolism by enhancing the collation of knowledge, interpretation of data, and prediction of metabolic capabilities. A frequent assumption in the use of genome-scale models is that the in vivo organism is evolved for optimal growth, where growth is represented by flux through a biomass objective function (BOF). While the specific composition of the BOF is crucial, its formulation is often inherited from similar organisms due to the experimental challenges associated with its proper determination. A cell's macro-molecular composition is not fixed and it responds to changes in environmental conditions. As a consequence, initiatives for the high-fidelity determination of cellular biomass composition have been launched. Thus, there is a need for a mathematical and computational framework capable of using multiple measurements of cellular biomass composition in different environments. Here, we propose two different computational approaches for directly addressing this challenge: Biomass Trade-off Weighting (BTW) and Higher-dimensional-plane InterPolation (HIP). In lieu of experimental data on biomass composition-variation in response to changing nutrient environment, we assess the properties of BTW and HIP using three hypothetical, yet biologically plausible, BOFs for the Escherichia coli genome-scale metabolic model iML1515. We find that the BTW and HIP formulations have a significant impact on model performance and phenotypes. Furthermore, the BTW method generates larger growth rates in all environments when compared to HIP. Using acetate secretion and the respiratory quotient as proxies for phenotypic changes, we find marked differences between the methods as HIP generates BOFs more similar to a reference BOF than BTW. We conclude that the presented methods constitute a conceptual step in developing genome-scale metabolic modelling approaches capable of addressing the inherent dependence of cellular biomass composition on nutrient environments. Christian Schulz 0011, Tjasa Kumelj, Emil Karlsen, Eivind Almaas |
PLoS Comput. Biol. | 4 |
| 2020 | ErrorTracer: an algorithm for identifying the origins of inconsistencies in genome-scale metabolic modelsabstractMOTIVATION: The number and complexity of genome-scale metabolic models is steadily increasing, empowered by automated model-generation algorithms. The quality control of the models, however, has always remained a significant challenge, the most fundamental being reactions incapable of carrying flux. Numerous automated gap-filling algorithms try to address this problem, but can rarely resolve all of a model's inconsistencies. The need for fast inconsistency checking algorithms has also been emphasized with the recent community push for automated model-validation before model publication. Previously, we wrote a graphical software to allow the modeller to solve the remaining errors manually. Nevertheless, model size and complexity remained a hindrance to efficiently tracking origins of inconsistency. RESULTS: We developed the ErrorTracer algorithm in order to address the shortcomings of existing approaches: ErrorTracer searches for inconsistencies, classifies them and identifies their origins. The algorithm is ∼2 orders of magnitude faster than current community standard methods, using only seconds even for large-scale models. This allows for interactive exploration in direct combination with model visualization, markedly simplifying the whole error-identification and correction work flow. AVAILABILITY AND IMPLEMENTATION: Windows and Linux executables and source code are available under the EPL 2.0 Licence at https://github.com/TheAngryFox/ModelExplorer and https://www.ntnu.edu/almaaslab/downloads. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Nikolay Martyushenko, Eivind Almaas |
Bioinform. | 2 |
| 2019 | ModelExplorer - software for visual inspection and inconsistency correction of genome-scale metabolic reconstructionsabstractBACKGROUND: Genome-scale metabolic network reconstructions are low level chemical representations of biological organisms. These models allow the system-level investigation of metabolic phenotypes using a variety of computational approaches. The link between a metabolic network model and an organisms' higher-level behaviour is usually found using a constraint-based analysis approach, such as FBA (Flux Balance Analysis). However, the process of model reconstruction rarely proceeds without error. Often, considerable parts of a model cannot carry flux under any condition. This is termed model inconsistency and is caused by faulty topology and/or stoichiometry of the underlying reconstructed network. While there exist several automated gap-filling tools that may solve some of the inconsistencies, much of the work still needs to be carried out manually. The common "linear list" format of writing biochemical reactions makes it difficult to intuit what is at the root of the inconsistent behaviour. Unfortunately, we have frequently observed that model builders do not correct their models past the abilities of automated tools, leaving many widely used models significantly inconsistent. RESULTS: We have developed the software ModelExplorer, which main purpose is to fill this gap by providing an intuitive and visual framework that allows the user to explore and correct inconsistencies in genome-scale metabolic models. The software will automatically visualize metabolic networks as graphs with distinct separation and delineation of cellular compartments. ModelExplorer highlights reactions and species that are unable to carry flux (blocked), with several different consistency checking modes available. Our software also allows the automatic identification of neighbours and production pathways of any species or reaction. Additionally, the user may focus on any chosen inconsistent part of the model on its own. This facilitates a rapid and visual identification of reactions and species responsible for model inconsistencies. Finally, ModelExplorer lets the user freely edit, add or delete model elements, allowing straight-forward correction of discovered issues. CONCLUSION: Overall, ModelExplorer is currently the fastest real-time metabolic network visualization program available. It implements several consistency checking algorithms, which in combination with its set of tracking tools, gives an efficient and systematic model-correction process. Nikolay Martyushenko, Eivind Almaas |
BMC Bioinform. | 2 |
| 2019 | Assessment of weighted topological overlap (wTO) to improve fidelity of gene co-expression networksabstractBACKGROUND: For more than a decade, gene expression data sets have been used as basis for the construction of co-expression networks used in systems biology investigations, leading to many important discoveries in a wide range of subjects spanning human disease to evolution and the development of organisms. A commonly encountered challenge in such investigations is first that of detecting, then subsequently removing, spurious correlations (i.e. links) in these networks. While access to a large number of measurements per gene would reduce this problem, often only a small number of measurements are available. The weighted Topological Overlap (wTO) measure, which incorporates information from the shared network-neighborhood of a given gene-pair into a single score, is a metric that is frequently used with the implicit expectation of producing higher-quality networks. However, the actual extent to which wTO improves on the accuracy of a co-expression analysis has not been quantified. RESULTS: Here, we used a large-sample biological data set containing 338 gene-expression measurements per gene as a reference system. From these data, we generated ensembles consisting of 10, 20 and 50 randomly selected measurements to emulate low-quality data sets, finding that the wTO measure consistently generates more robust scores than what results from simple correlation calculations. Furthermore, for the data sets consisting of only 10 and 20 samples per gene, we find that wTO serves as a better predictor of the correlation scores generated from the full data set. However, we find that using wTO as a score for network building substantially alters several topographical aspects of the resulting networks, with no conclusive evidence that the resulting structure is more accurate. Importantly, we find that the much used approach of applying a soft-threshold modifier to link weights prior to computing the wTO substantially decreases the robustness of the resulting wTO network, but increases the predictive power of wTO networks with regards to the reference correlation (soft threshold) network, particularly as the size of the data sets increases. CONCLUSION: Our analysis demonstrates that, in agreement with previous assumptions, the wTO approach is capable of significantly improving the fidelity of co-expression networks, and that this effect is especially evident for cases of low-sample number gene-expression data sets. André Voigt, Eivind Almaas |
BMC Bioinform. | 2 |
| 2018 | wTO: an R package for computing weighted topological overlap and a consensus network with integrated visualization toolabstractBACKGROUND: Network analyses, such as of gene co-expression networks, metabolic networks and ecological networks have become a central approach for the systems-level study of biological data. Several software packages exist for generating and analyzing such networks, either from correlation scores or the absolute value of a transformed score called weighted topological overlap (wTO). However, since gene regulatory processes can up- or down-regulate genes, it is of great interest to explicitly consider both positive and negative correlations when constructing a gene co-expression network. RESULTS: Here, we present an R package for calculating the weighted topological overlap (wTO), that, in contrast to existing packages, explicitly addresses the sign of the wTO values, and is thus especially valuable for the analysis of gene regulatory networks. The package includes the calculation of p-values (raw and adjusted) for each pairwise gene score. Our package also allows the calculation of networks from time series (without replicates). Since networks from independent datasets (biological repeats or related studies) are not the same due to technical and biological noise in the data, we additionally, incorporated a novel method for calculating a consensus network (CN) from two or more networks into our R package. To graphically inspect the resulting networks, the R package contains a visualization tool, which allows for the direct network manipulation and access of node and link information. When testing the package on a standard laptop computer, we can conduct all calculations for systems of more than 20,000 genes in under two hours. We compare our new wTO package to state of art packages and demonstrate the application of the wTO and CN functions using 3 independently derived datasets from healthy human pre-frontal cortex samples. To showcase an example for the time series application we utilized a metagenomics data set. CONCLUSION: In this work, we developed a software package that allows the computation of wTO networks, CNs and a visualization tool in the R statistical environment. It is publicly available on CRAN repositories under the GPL -2 Open Source License ( https://cran.r-project.org/web/packages/wTO/ ). Deisy Morselli Gysi, André Voigt, Tiago de Miranda Fragoso, Eivind Almaas, Katja Nowick |
BMC Bioinform. | 4 |
| 2018 | Automated generation of genome-scale metabolic draft reconstructions based on KEGGabstractBACKGROUND: Constraint-based modeling is a widely used and powerful methodology to assess the metabolic phenotypes and capabilities of an organism. The starting point and cornerstone of all such modeling is a genome-scale metabolic network reconstruction. The creation, further development, and application of such networks is a growing field of research thanks to a plethora of readily accessible computational tools. While the majority of studies are focused on single-species analyses, typically of a microbe, the computational study of communities of organisms is gaining attention. Similarly, reconstructions that are unified for a multi-cellular organism have gained in popularity. Consequently, the rapid generation of genome-scale metabolic reconstructed networks is crucial. While multiple web-based or stand-alone tools are available for automated network reconstruction, there is, however, currently no publicly available tool that allows the swift assembly of draft reconstructions of community metabolic networks and consolidated metabolic networks for a specified list of organisms. RESULTS: Here, we present AutoKEGGRec, an automated tool that creates first draft metabolic network reconstructions of single organisms, community reconstructions based on a list of organisms, and finally a consolidated reconstruction for a list of organisms or strains. AutoKEGGRec is developed in Matlab and works seamlessly with the COBRA Toolbox v3, and it is based on only using the KEGG database as external input. The generated first draft reconstructions are stored in SBML files and consist of all reactions for a KEGG organism ID and corresponding linked genes. This provides a comprehensive starting point for further refinement and curation using the host of COBRA toolbox functions or other preferred tools. Through the data structures created, the tool also facilitates a comparative analysis of metabolic content in any given number of organisms present in the KEGG database. CONCLUSION: AutoKEGGRec provides a first step in a metabolic network reconstruction process, filling a gap for tools creating community and consolidated metabolic networks. Based only on KEGG data as external input, the generated reconstructions consist of data with a directly traceable foundation and pedigree. With AutoKEGGRec, this kind of modeling is made accessible to a wider part of the genome-scale metabolic analysis community. Emil Karlsen, Christian Schulz 0011, Eivind Almaas |
BMC Bioinform. | 3 |
| 2017 | A composite network of conserved and tissue specific gene interactions reveals possible genetic interactions in gliomaabstractDifferential co-expression network analyses have recently become an important step in the investigation of cellular differentiation and dysfunctional gene-regulation in cell and tissue disease-states. The resulting networks have been analyzed to identify and understand pathways associated with disorders, or to infer molecular interactions. However, existing methods for differential co-expression network analysis are unable to distinguish between various forms of differential co-expression. To close this gap, here we define the three different kinds (conserved, specific, and differentiated) of differential co-expression and present a systematic framework, CSD, for differential co-expression network analysis that incorporates these interactions on an equal footing. In addition, our method includes a subsampling strategy to estimate the variance of co-expressions. Our framework is applicable to a wide variety of cases, such as the study of differential co-expression networks between healthy and disease states, before and after treatments, or between species. Applying the CSD approach to a published gene-expression data set of cerebral cortex and basal ganglia samples from healthy individuals, we find that the resulting CSD network is enriched in genes associated with cognitive function, signaling pathways involving compounds with well-known roles in the central nervous system, as well as certain neurological diseases. From the CSD analysis, we identify a set of prominent hubs of differential co-expression, whose neighborhood contains a substantial number of genes associated with glioblastoma. The resulting gene-sets identified by our CSD analysis also contain many genes that so far have not been recognized as having a role in glioblastoma, but are good candidates for further studies. CSD may thus aid in hypothesis-generation for functional disease-associations. André Voigt, Katja Nowick, Eivind Almaas |
PLoS Comput. Biol. | 3 |
| 2007 | Trend Motif: A Graph Mining Approach for Analysis of Dynamic Complex NetworksabstractComplex networks have been used successfully in scientific disciplines ranging from sociology to microbiology to describe systems of interacting units. Until recently, studies of complex networks have mainly focused on their network topology. However, in many real world applications, the edges and vertices have associated attributes that are frequently represented as vertex or edge weights. Furthermore, these weights are often not static, instead changing with time and forming a time series. Hence, to fully understand the dynamics of the complex network, we have to consider both network topology and related time series data. In this work, we propose a motif mining approach to identify trend motifs for such purposes. Simply stated, a trend motif describes a recurring subgraph where each of its vertices or edges displays similar dynamics over a user- defined period. Given this, each trend motif occurrence can help reveal significant events in a complex system; frequent trend motifs may aid in uncovering dynamic rules of change for the system, and the distribution of trend motifs may characterize the global dynamics of the system. Here, we have developed efficient mining algorithms to extract trend motifs. Our experimental validation using three disparate empirical datasets, ranging from the stock market, world trade, to a protein interaction network, has demonstrated the efficiency and effectiveness of our approach. Ruoming Jin, Scott McCallen, Eivind Almaas |
ICDM | 3 |
| 2005 | The Activity Reaction Core and Plasticity of Metabolic NetworksabstractUnderstanding the system-level adaptive changes taking place in an organism in response to variations in the environment is a key issue of contemporary biology. Current modeling approaches, such as constraint-based flux-balance analysis, have proved highly successful in analyzing the capabilities of cellular metabolism, including its capacity to predict deletion phenotypes, the ability to calculate the relative flux values of metabolic reactions, and the capability to identify properties of optimal growth states. Here, we use flux-balance analysis to thoroughly assess the activity of Escherichia coli, Helicobacter pylori, and Saccharomyces cerevisiae metabolism in 30,000 diverse simulated environments. We identify a set of metabolic reactions forming a connected metabolic core that carry non-zero fluxes under all growth conditions, and whose flux variations are highly correlated. Furthermore, we find that the enzymes catalyzing the core reactions display a considerably higher fraction of phenotypic essentiality and evolutionary conservation than those catalyzing noncore reactions. Cellular metabolism is characterized by a large number of species-specific conditionally active reactions organized around an evolutionary conserved, but always active, metabolic core. Finally, we find that most current antibiotics interfering with bacterial metabolism target the core enzymes, indicating that our findings may have important implications for antimicrobial drug-target discovery. Eivind Almaas, Zoltán N. Oltvai, Albert-László Barabási |
PLoS Comput. Biol. | 1 |