EDBT 2026 Demo / reviewers in the wild / expert
Andreas Hellander
dblp:19/1432
· DBLP profile ↗
27ranked-venue papers
0as first author
15since 2021 · last 2025
0000-0001-7273-7923ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 17 · 8 since 2021Artificial intelligence and machine learning · 5 · 4 since 2021Systems, architecture and hardware · 3 · 1 since 2021Software engineering, systems software and programming languages · 3Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Accelerating Simulation-Based Inference with Variational AutoencodersabstractWe present a generative modeling approach based on the variational inference framework for likelihood-free simulation-based inference. The method leverages latent variables within variational autoencoders to efficiently estimate complex posterior distributions arising from stochastic simulations. We explore two variations of this approach distinguished by their treatment of the prior distribution. The first model adapts the prior based on observed data using a multivariate prior network, enhancing generalization across various posterior queries. In contrast, the second model utilizes a standard Gaussian prior, offering simplicity while still effectively capturing complex posterior distributions. We demonstrate the ability of the proposed approach to approximate complex posteriors while maintaining computational efficiency on well-established benchmark problems. Mayank Nautiyal, Andrey Shternshis, Andreas Hellander |
IJCNN | 3 |
| 2025 | Exploiting the Asymmetric Uncertainty Structure of Pre-trained VLMs on the Unit HypersphereabstractVision-language models (VLMs) as foundation models have significantly enhanced performance across a wide range of visual and textual tasks, without requiring large-scale training from scratch for downstream tasks. However, these deterministic VLMs fail to capture the inherent ambiguity and uncertainty in natural language and visual data. Recent probabilistic post-hoc adaptation methods address this by mapping deterministic embeddings onto probability distributions; however, existing approaches do not account for the asymmetric uncertainty between modalities, and the constraint that meaningful deterministic embeddings reside on a unit hypersphere, potentially leading to suboptimal performance. In this paper, we address the asymmetric uncertainty structure inherent in textual and visual data, and propose AsymVLM to build probabilistic embeddings from pre-trained VLMs on the unit hypersphere, enabling uncertainty quantification. We validate the effectiveness of the probabilistic embeddings on established benchmarks, and present comprehensive ablation studies demonstrating the inherent nature of asymmetry in the uncertainty structure of textual and visual data. Max Andersson, Stina Fredriksson, Edward Glöckner, Andreas Hellander, Ekta Vats |
NeurIPS | 5 |
| 2024 | Data management of scientific applications in a reinforcement learning-based hierarchical storage systemabstractIn many areas of data-driven science, large datasets are generated where the individual data objects are images, matrices, or otherwise have a clear structure. However, these objects can be information-sparse, and a challenge is to efficiently find and work with the most interesting data as early as possible in an analysis pipeline. We have recently proposed a new model for big data management where the internal structure and information of the data are associated with each data object (as opposed to simple metadata). There is then an opportunity for comprehensive data management solutions to account for data-specific internal structure as well as access patterns. In this article, we explore this idea together with our recently proposed hierarchical storage management framework that uses reinforcement learning (RL) for autonomous and dynamic data placement in different tiers in a storage hierarchy. Our case-study is based on four scientific datasets: Protein translocation microscopy images, Airfoil angle of attack meshes, 1000 Genomes sequences, and Phenotypic screening images. The presented results highlight that our framework is optimal and can quickly adapt to new data access requirements. It overall reduces the data processing time, and the proposed autonomous data placement is superior compared to any static or semi-static data placement policies. Tianru Zhang, Ankit Gupta 0018, María Andreína Francisco Rodríguez, Ola Spjuth, Andreas Hellander, Salman Zubair Toor |
Expert Syst. Appl. | 5 |
| 2023 | Efficient Hierarchical Storage Management Empowered by Reinforcement Learning Extended AbstractabstractWith the rapid development of big data and cloud computing, data management has become increasingly challenging. A possible solution is to use an intelligent hierarchical (multi-tier) storage system (HSS). An HSS is a meta solution that consists of different storage frameworks organized as a jointly constructed storage pool. A built-in data migration policy that determines the optimal placement of the datasets in the hierarchy is essential. Placement decisions are a non-trivial task since they should be made according to the characteristics of the dataset, the tier status in a hierarchy, and access patterns. This paper presents an open-source hierarchical storage framework with a dynamic migration policy based on reinforcement learning (RL). Tianru Zhang, Andreas Hellander, Salman Zubair Toor |
ICDE | 2 |
| 2023 | Contributions of cell behavior to geometric order in embryonic cartilageabstractDuring early development, cartilage provides shape and stability to the embryo while serving as a precursor for the skeleton. Correct formation of embryonic cartilage is hence essential for healthy development. In vertebrate cranial cartilage, it has been observed that a flat and laterally extended macroscopic geometry is linked to regular microscopic structure consisting of tightly packed, short, transversal clonar columns. However, it remains an ongoing challenge to identify how individual cells coordinate to successfully shape the tissue, and more precisely which mechanical interactions and cell behaviors contribute to the generation and maintenance of this columnar cartilage geometry during embryogenesis. Here, we apply a three-dimensional cell-based computational model to investigate mechanical principles contributing to column formation. The model accounts for clonal expansion, anisotropic proliferation and the geometrical arrangement of progenitor cells in space. We confirm that oriented cell divisions and repulsive mechanical interactions between cells are key drivers of column formation. In addition, the model suggests that column formation benefits from the spatial gaps created by the extracellular matrix in the initial configuration, and that column maintenance is facilitated by sequential proliferative phases. Our model thus correctly predicts the dependence of local order on division orientation and tissue thickness. The present study presents the first cell-based simulations of cell mechanics during cranial cartilage formation and we anticipate that it will be useful in future studies on the formation and growth of other cartilage geometries. Sonja Mathias, Igor Adameyko, Andreas Hellander, Jochen Kursawe |
PLoS Comput. Biol. | 3 |
| 2023 | Efficient Hierarchical Storage Management Empowered by Reinforcement LearningabstractWith the rapid development of big data and cloud computing, data management has become increasingly challenging. Over the years, a number of frameworks for data management have become available. Most of them are highly efficient, but ultimately create data silos. It becomes difficult to move and work coherently with data as new requirements emerge. A possible solution is to use an intelligent hierarchical (multi-tier) storage system (HSS). A HSS is a meta solution that consists of different storage frameworks organized as a jointly constructed storage pool. A built-in data migration policy that determines the optimal placement of the datasets in the hierarchy is essential. Placement decisions is a non-trivial task since it should be made according to the characteristics of the dataset, the tier status in a hierarchy, and access patterns. This paper presents an open-source hierarchical storage framework with a dynamic migration policy based on reinforcement learning (RL). We present a mathematical model, a software architecture, and implementations based on both simulations and a live cloud-based environment. We compare the proposed RL-based strategy to a baseline of three rule-based policies, showing that the RL-based policy achieves significantly higher efficiency and optimal data distribution in different scenarios. Tianru Zhang, Andreas Hellander, Salman Zubair Toor |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Scalable federated machine learning with FEDnabstractFederated machine learning promises to overcome the input privacy challenge in machine learning. By iteratively updating a model on private clients and aggregating these local model updates into a global federated model, private data is incorporated in the federated model without needing to share and expose that data. Several open software projects for federated learning have appeared. Most of them focuses on supporting flexible experimentation with different model aggregation schemes and with different privacy-enhancing technologies. However, there is a lack of open frameworks that focuses on critical distributed computing aspects of the problem such as scalability and resilience. It is a big step to take for a data scientist to go from an experimental sandbox to testing their federated schemes at scale in real-world geographically distributed settings. To bridge this gap we have designed and developed a production-grade hierarchical federated learning framework, FEDn. The framework is specifically designed to make it easy to go from local development in pseudo-distributed mode to horizontally scalable distributed deployments. FEDn both aims to be production grade for industrial applications and a flexible research tool to explore real-world performance of novel federated algorithms and the framework has been used in number of industrial and academic R&D projects. In this paper we present the architecture and implementation of FEDn. We demonstrate the framework's scalability and efficiency in evaluations based on two case-studies representative for a cross-silo and a cross-device use-case respectively. Morgan Ekmefjord, Addi Ait-Mlouk, Sadi Alawadi, Mattias Åkesson, Ola Spjuth, Salman Zubair Toor, Andreas Hellander |
CCGRID | 8 |
| 2022 | Robust and integrative Bayesian neural networks for likelihood-free parameter inferenceabstractState-of-the-art neural network-based methods for learning summary statistics have delivered promising results for simulation-based likelihood-free parameter inference. Existing approaches for learning summarizing networks are mainly based on deterministic neural networks, and do not take network prediction uncertainty into account. This work proposes a robust integrated approach that learns summary statistics using Bayesian neural networks, and produces a proposal posterior density using categorical distributions. An adaptive sampling scheme selects simulation locations to efficiently and iteratively refine the predictive proposal posterior of the network conditioned on observations. This allows for more efficient and robust convergence on comparatively large prior spaces. The approximated proposal posterior can then either be processed through a correction mechanism, or be used in conjunction with a density estimator to arrive at the true posterior. We demonstrate our approach on benchmark examples. Fredrik Wrede, Robin Eriksson, Richard M. Jiang 0002, Linda R. Petzold, Stefan Engblom, Andreas Hellander |
IJCNN | 6 |
| 2022 | CBMOS: a GPU-enabled Python framework for the numerical study of center-based modelsabstractBACKGROUND: Cell-based models are becoming increasingly popular for applications in developmental biology. However, the impact of numerical choices on the accuracy and efficiency of the simulation of these models is rarely meticulously tested. Without concrete studies to differentiate between solid model conclusions and numerical artifacts, modelers are at risk of being misled by their experiments' results. Most cell-based modeling frameworks offer a feature-rich environment, providing a wide range of biological components, but are less suitable for numerical studies. There is thus a need for software specifically targeted at this use case. RESULTS: We present CBMOS, a Python framework for the simulation of the center-based or cell-centered model. Contrary to other implementations, CBMOS' focus is on facilitating numerical study of center-based models by providing access to multiple ordinary differential equation solvers and force functions through a flexible, user-friendly interface and by enabling rapid testing through graphics processing unit (GPU) acceleration. We show-case its potential by illustrating two common workflows: (1) comparison of the numerical properties of two solvers within a Jupyter notebook and (2) measuring average wall times of both solvers on a high performance computing cluster. More specifically, we confirm that although for moderate accuracy levels the backward Euler method allows for larger time step sizes than the commonly used forward Euler method, its additional computational cost due to being an implicit method prohibits its use for practical test cases. CONCLUSIONS: CBMOS is a flexible, easy-to-use Python implementation of the center-based model, exposing both basic model assumptions and numerical components to the user. It is available on GitHub and PyPI under an MIT license. CBMOS allows for fast prototyping on a central processing unit for small systems through the use of NumPy. Using CuPy on a GPU, cell populations of up to 10,000 cells can be simulated within a few seconds. As such, it will substantially lower the time investment for any modeler to check the crucial assumption that model conclusions are independent of numerical issues. Sonja Mathias, Adrien Coulier, Andreas Hellander |
BMC Bioinform. | 3 |
| 2022 | Systematic comparison of modeling fidelity levels and parameter inference settings applied to negative feedback gene regulationabstractQuantitative stochastic models of gene regulatory networks are important tools for studying cellular regulation. Such models can be formulated at many different levels of fidelity. A practical challenge is to determine what model fidelity to use in order to get accurate and representative results. The choice is important, because models of successively higher fidelity come at a rapidly increasing computational cost. In some situations, the level of detail is clearly motivated by the question under study. In many situations however, many model options could qualitatively agree with available data, depending on the amount of data and the nature of the observations. Here, an important distinction is whether we are interested in inferring the true (but unknown) physical parameters of the model or if it is sufficient to be able to capture and explain available data. The situation becomes complicated from a computational perspective because inference needs to be approximate. Most often it is based on likelihood-free Approximate Bayesian Computation (ABC) and here determining which summary statistics to use, as well as how much data is needed to reach the desired level of accuracy, are difficult tasks. Ultimately, all of these aspects-the model fidelity, the available data, and the numerical choices for inference-interplay in a complex manner. In this paper we develop a computational pipeline designed to systematically evaluate inference accuracy for a wide range of true known parameters. We then use it to explore inference settings for negative feedback gene regulation. In particular, we compare a detailed spatial stochastic model, a coarse-grained compartment-based multiscale model, and the standard well-mixed model, across several data-scenarios and for multiple numerical options for parameter inference. Practically speaking, this pipeline can be used as a preliminary step to guide modelers prior to gathering experimental data. By training Gaussian processes to approximate the distance function values, we are able to substantially reduce the computational cost of running the pipeline. Adrien Coulier, Marc Sturrock, Andreas Hellander |
PLoS Comput. Biol. | 4 |
| 2022 | Identification of dynamic mass-action biochemical reaction networks using sparse Bayesian methodsabstractIdentifying the reactions that govern a dynamical biological system is a crucial but challenging task in systems biology. In this work, we present a data-driven method to infer the underlying biochemical reaction system governing a set of observed species concentrations over time. We formulate the problem as a regression over a large, but limited, mass-action constrained reaction space and utilize sparse Bayesian inference via the regularized horseshoe prior to produce robust, interpretable biochemical reaction networks, along with uncertainty estimates of parameters. The resulting systems of chemical reactions and posteriors inform the biologist of potentially several reaction systems that can be further investigated. We demonstrate the method on two examples of recovering the dynamics of an unknown reaction system, to illustrate the benefits of improved accuracy and information obtained. Richard M. Jiang 0002, Fredrik Wrede, Andreas Hellander, Linda R. Petzold |
PLoS Comput. Biol. | 4 |
| 2022 | Convolutional Neural Networks as Summary Statistics for Approximate Bayesian ComputationabstractApproximate Bayesian Computation is widely used in systems biology for inferring parameters in stochastic gene regulatory network models. Its performance hinges critically on the ability to summarize high-dimensional system responses such as time series into a few informative, low-dimensional summary statistics. The quality of those statistics acutely impacts the accuracy of the inference task. Existing methods to select the best subset out of a pool of candidate statistics do not scale well with large pools of several tens to hundreds of candidate statistics. Since high quality statistics are imperative for good performance, this becomes a serious bottleneck when performing inference on complex and high-dimensional problems. This paper proposes a convolutional neural network architecture for automatically learning informative summary statistics of temporal responses. We show that the proposed network can effectively circumvent the statistics selection problem of the preprocessing step for ABC inference. The proposed approach is demonstrated on two benchmark problem and one challenging inference problem learning parameters in a high-dimensional stochastic genetic oscillator. We also study the impact of experimental design on network performance by comparing different data richness and data acquisition strategies. Mattias Åkesson, Fredrik Wrede, Andreas Hellander |
IEEE ACM Trans. Comput. Biol. Bioinform. | 4 |
| 2021 | Epidemiological modeling in StochSS Live!abstractSUMMARY: We present StochSS Live!, a web-based service for modeling, simulation and analysis of a wide range of mathematical, biological and biochemical systems. Using an epidemiological model of COVID-19, we demonstrate the power of StochSS Live! to enable researchers to quickly develop a deterministic or a discrete stochastic model, infer its parameters and analyze the results. AVAILABILITY AND IMPLEMENTATION: StochSS Live! is freely available at https://live.stochss.org/. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Richard M. Jiang 0002, Bruno Jacob, Matthew Geiger, Sean Matthew, Bryan Rumsey, Fredrik Wrede, Tau-Mu Yi, Brian Drawert, Andreas Hellander, Linda R. Petzold |
Bioinform. | 10 |
| 2021 | Scalable machine learning-assisted model exploration and inference using SciopeabstractSUMMARY: Discrete stochastic models of gene regulatory networks are fundamental tools for in silico study of stochastic gene regulatory networks. Likelihood-free inference and model exploration are critical applications to study a system using such models. However, the massive computational cost of complex, high-dimensional and stochastic modelling currently limits systematic investigation to relatively simple systems. Recently, machine-learning-assisted methods have shown great promise to handle larger, more complex models. To support both ease-of-use of this new class of methods, as well as their further development, we have developed the scalable inference, optimization and parameter exploration (Sciope) toolbox. Sciope is designed to support new algorithms for machine-learning-assisted model exploration and likelihood-free inference. Moreover, it is built ground up to easily leverage distributed and heterogeneous computational resources for convenient parallelism across platforms from workstations to clouds. AVAILABILITY AND IMPLEMENTATION: The Sciope Python3 toolbox is freely available on https://github.com/Sciope/Sciope, and has been tested on Linux, Windows and macOS platforms. SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online. Fredrik Wrede, Andreas Hellander |
Bioinform. | 3 |
| 2021 | Accelerated regression-based summary statistics for discrete stochastic systems via approximate simulatorsabstractBACKGROUND: Approximate Bayesian Computation (ABC) has become a key tool for calibrating the parameters of discrete stochastic biochemical models. For higher dimensional models and data, its performance is strongly dependent on having a representative set of summary statistics. While regression-based methods have been demonstrated to allow for the automatic construction of effective summary statistics, their reliance on first simulating a large training set creates a significant overhead when applying these methods to discrete stochastic models for which simulation is relatively expensive. In this τ work, we present a method to reduce this computational burden by leveraging approximate simulators of these systems, such as ordinary differential equations and τ-Leaping approximations. RESULTS: We have developed an algorithm to accelerate the construction of regression-based summary statistics for Approximate Bayesian Computation by selectively using the faster approximate algorithms for simulations. By posing the problem as one of ratio estimation, we use state-of-the-art methods in machine learning to show that, in many cases, our algorithm can significantly reduce the number of simulations from the full resolution model at a minimal cost to accuracy and little additional tuning from the user. We demonstrate the usefulness and robustness of our method with four different experiments. CONCLUSIONS: We provide a novel algorithm for accelerating the construction of summary statistics for stochastic biochemical systems. Compared to the standard practice of exclusively training from exact simulator samples, our method is able to dramatically reduce the number of required calls to the stochastic simulator at a minimal loss in accuracy. This can immediately be implemented to increase the overall speed of the ABC workflow for estimating parameters in complex systems. Richard M. Jiang 0002, Fredrik Wrede, Andreas Hellander, Linda R. Petzold |
BMC Bioinform. | 4 |
| 2020 | Smart Resource Management for Data Streaming using an Online Bin-packing StrategyabstractData stream processing frameworks provide reliable and efficient mechanisms for executing complex workflows over large datasets. A common challenge for the majority of currently available streaming frameworks is efficient utilization of resources. Most frameworks use static or semi-static settings for resource utilization that work well for established use cases but lead to marginal improvements for unseen scenarios. Another pressing issue is the efficient processing of large individual objects such as images and matrices typical for scientific datasets. HarmonicIO has proven to be a good solution for streams of relatively large individual objects, as demonstrated in a benchmark comparison with the Apache Spark and Kafka streaming frameworks. We here present an extension of the HarmonicIO framework based on the online bin-packing algorithm. The main focus is to compare different strategies adapted in streaming frameworks for efficient resource utilization. Based on a real world use case from large-scale microscopy pipelines, we compare two different strategies of auto-scaling implemented in the HarmonicIO and Spark Streaming frameworks. Oliver Stein, Ben Blamey, Alan Sabirsh, Ola Spjuth, Andreas Hellander, Salman Zubair Toor |
IEEE BigData | 6 |
| 2019 | Adapting the Secretary Hiring Problem for Optimal Hot-Cold Tier Placement Under Top-K WorkloadsabstractTop-K queries are an established heuristic in information retrieval. This paper presents an approach for optimal tiered storage allocation under stream processing workloads using this heuristic: those requiring the analysis of only the top-K ranked most relevant documents from a fixed-length stream, stream window, or batch job. Documents are ranked for relevance on a user-specified interestingness function, the top-K stored for further processing. This scenario bears similarity to the classic Secretary Hiring Problem (SHP), and the expected rate of document writes and document lifetime can be modelled as a function of document index. We present parameter-based algorithms for storage tier placement, minimizing document storage and transport costs. We derive expressions for optimal parameter values in terms of tier storage and transport costs a priori, without needing to monitor the application. This contrasts with (often complex) existing work on tiered storage optimization, which is either tightly coupled to specific use cases, or requires active monitoring of application IO load - ill-suited to long-running or one-off operations common in the scientific computing domain. We motivate and evaluate our model with a trace-driven simulation of human-in-the-loop bio-chemical model exploration, and two cloud storage case studies. Ben Blamey, Fredrik Wrede, Andreas Hellander, Salman Zubair Toor |
CCGRID | 4 |
| 2019 | Smart computational exploration of stochastic gene regulatory network models using human-in-the-loop semi-supervised learningabstractMOTIVATION: Discrete stochastic models of gene regulatory network models are indispensable tools for biological inquiry since they allow the modeler to predict how molecular interactions give rise to nonlinear system output. Model exploration with the objective of generating qualitative hypotheses about the workings of a pathway is usually the first step in the modeling process. It involves simulating the gene network model under a very large range of conditions, due to the large uncertainty in interactions and kinetic parameters. This makes model exploration highly computational demanding. Furthermore, with no prior information about the model behavior, labor-intensive manual inspection of very large amounts of simulation results becomes necessary. This limits systematic computational exploration to simplistic models. RESULTS: We have developed an interactive, smart workflow for model exploration based on semi-supervised learning and human-in-the-loop labeling of data. The workflow lets a modeler rapidly discover ranges of interesting behaviors predicted by the model. Utilizing that similar simulation output is in proximity of each other in a feature space, the modeler can focus on informing the system about what behaviors are more interesting than others by labeling, rather than analyzing simulation results with custom scripts and workflows. This results in a large reduction in time-consuming manual work by the modeler early in a modeling project, which can substantially reduce the time needed to go from an initial model to testable predictions and downstream analysis. AVAILABILITY AND IMPLEMENTATION: A python-package is available at https://github.com/Wrede/mio.git. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Fredrik Wrede, Andreas Hellander |
Bioinform. | 2 |
| 2018 | HarmonicIO: Scalable Data Stream Processing for Scientific DatasetsabstractMany streaming frameworks have been introduced to deal with the needs for online analysis of massive datasets. Scientific applications often require significant changes to make them compatible with these frameworks. Other issues include tight coupling with the underlying infrastructure, shared computing environment, static topology settings, and complex configuration. In this article we present HarmonicIO, a lightweight streaming framework specialized for scientific datasets. It boasts a smart dynamic architecture, is highly elastic, and enforces a clear separation between framework components and application execution environment using container technology. Preechakorn Torruangwatthana, Håkan Wieslander, Ben Blamey, Andreas Hellander, Salman Zubair Toor |
IEEE CLOUD | 4 |
| 2018 | Orchestral: A Lightweight Framework for Parallel Simulations of Cell-Cell CommunicationabstractWe develop a modeling and simulation framework capable of massively parallel simulation of multicellular systems with spatially resolved stochastic kinetics in individual cells. By the use of operator-splitting we decouple the simulation of reaction-diffusion kinetics inside the cells from the simulation of molecular cell-cell interactions occurring on the boundaries between cells. This decoupling leverages the inherent scale separation in the underlying model to enable highly horizontally scalable parallel simulation, suitable for simulation on heterogeneous, distributed computing infrastructures such as public and private clouds. Thanks to its modular structure, our frameworks makes it possible to couple just any existing single-cell simulation software together with any cell signaling simulator. We exemplify the flexibility and scalability of the framework by using the popular single-cell simulation software eGFRD to construct and simulate a multicellular model of Notch-Delta signaling over OpenStack cloud infrastructure provided by the SNIC Science Cloud. Adrien Coulier, Andreas Hellander |
eScience | 2 |
| 2018 | BAMSI: a multi-cloud service for scalable distributed filtering of massive genome dataabstractBACKGROUND: The advent of next-generation sequencing (NGS) has made whole-genome sequencing of cohorts of individuals a reality. Primary datasets of raw or aligned reads of this sort can get very large. For scientific questions where curated called variants are not sufficient, the sheer size of the datasets makes analysis prohibitively expensive. In order to make re-analysis of such data feasible without the need to have access to a large-scale computing facility, we have developed a highly scalable, storage-agnostic framework, an associated API and an easy-to-use web user interface to execute custom filters on large genomic datasets. RESULTS: We present BAMSI, a Software as-a Service (SaaS) solution for filtering of the 1000 Genomes phase 3 set of aligned reads, with the possibility of extension and customization to other sets of files. Unique to our solution is the capability of simultaneously utilizing many different mirrors of the data to increase the speed of the analysis. In particular, if the data is available in private or public clouds - an increasingly common scenario for both academic and commercial cloud providers - our framework allows for seamless deployment of filtering workers close to data. We show results indicating that such a setup improves the horizontal scalability of the system, and present a possible use case of the framework by performing an analysis of structural variation in the 1000 Genomes data set. CONCLUSIONS: BAMSI constitutes a framework for efficient filtering of large genomic data sets that is flexible in the use of compute as well as storage resources. The data resulting from the filter is assumed to be greatly reduced in size, and can easily be downloaded or routed into e.g. a Hadoop cluster for subsequent interactive analysis using Hive, Spark or similar tools. In this respect, our framework also suggests a general model for making very large datasets of high scientific value more accessible by offering the possibility for organizations to share the cost of hosting data on hot storage, without compromising the scalability of downstream analysis. Kristiina Ausmees, Aji John, Salman Zubair Toor, Andreas Hellander, Carl Nettelblad |
BMC Bioinform. | 4 |
| 2017 | Cost-aware Application Development and Management using CLOUD-METRIC
Alieu Jallow, Andreas Hellander, Salman Zubair Toor |
CLOSER | 2 |
| 2017 | SNIC Science Cloud (SSC): A National-Scale Cloud Infrastructure for Swedish AcademiaabstractThe cloud computing paradigm have fundamentally changed the way computational resources are being offered. Although the number of large-scale providers in academia is still relatively small, there is a rapidly increasing interest and adoption of cloud Infrastructure-as-a-Service in the scientific community. The added flexibility in how applications can be implemented compared to traditional batch computing systems is one of the key success factors for the paradigm, and scientific cloud computing promises to increase adoption of simulation and data analysis in scientific communities not traditionally users of large scale e-Infrastructure, the so called ”long tail of science”. In 2014, the Swedish National Infrastructure for Computing (SNIC) initiated a project to investigate the cost and constraints of offering cloud infrastructure for Swedish academia. The aim was to build a platform where academics could evaluate cloud computing for their use-cases. SNIC Science Cloud (SSC) has since then evolved into a national-scale cloud infrastructure based on three geographically distributed regions. In this article we present the SSC vision, architectural details and user stories. We summarize the experiences gained from running a nationalscale cloud facility into ”ten simple rules” for starting up a science cloud project based on OpenStack. We also highlight some key areas that require careful attention in order to offer cloud infrastructure for ubiquitous academic needs and in particular scientific workloads. Salman Zubair Toor, Mathias Lindberg, Ingemar Falman, Andreas Vallin, Olof Mohill, Pontus Freyhult, Linus Nilsson, Martin Agback, Lars Viklund, Henric Zazzik, Ola Spjuth, Marco Capuccini, Joakim Moller, Donal Murtagh, Andreas Hellander |
eScience | 15 |
| 2016 | Stochastic Simulation Service: Bridging the Gap between the Computational Expert and the BiologistabstractWe present StochSS: Stochastic Simulation as a Service, an integrated development environment for modeling and simulation of both deterministic and discrete stochastic biochemical systems in up to three dimensions. An easy to use graphical user interface enables researchers to quickly develop and simulate a biological model on a desktop or laptop, which can then be expanded to incorporate increasing levels of complexity. StochSS features state-of-the-art simulation engines. As the demand for computational power increases, StochSS can seamlessly scale computing resources in the cloud. In addition, StochSS can be deployed as a multi-user software environment where collaborators share computational resources and exchange models via a public model repository. We demonstrate the capabilities and ease of use of StochSS with an example of model development and simulation at increasing levels of complexity. Brian Drawert, Andreas Hellander, Benjamin B. Bales, Debjani Banerjee, Giovanni Bellesia, Bernie J. Daigle Jr., Geoffrey Douglas, Mengyuan Gu, Anand Gupta, Stefan Hellander, Christopher B. Horuk, Dibyendu Nath, Aviral Takkar, Sheng Wu 0002, Per Lötstedt, Chandra Krintz, Linda R. Petzold |
PLoS Comput. Biol. | 2 |
| 2013 | Scientific Analysis by Queries in Extended SPARQL over a Scalable e-Science Data StoreabstractData-intensive applications in e-Science require scalable solutions for storage as well as interactive tools for analysis of scientific data. It is important to be able to query the data in a storage-independent way, and to be able to obtain the results of the data-analysis incrementally (in contrast to traditional batch solutions). We use the RDF data model extended with multidimensional numeric arrays to represent the results, parameters, and other metadata describing scientific experiments, and SciSPARQL, an extension of the SPARQL language, to combine massive numeric array data and metadata in queries. To address the scalability problem we present an architecture that enables the same SciSPARQL queries to be executed on the RDF dataset whether it is stored in a relational DBMS or mapped over a specialized geographically distributed e-Science data store. In order to minimize access and communication costs, we represent the arrays with proxy objects, and retrieve their content lazily. We formulate typical analysis tasks from a computational biology application in terms of SciSPARQL queries, and compare the query processing performance with manually written scripts in MATLAB. Andrej Andrejev, Salman Zubair Toor, Andreas Hellander, Sverker Holmgren, Tore Risch |
e-Science | 3 |
| 2012 | Reducing Complexity in Management of eScience ComputationsabstractIn this paper we address reduction of complexity in management of scientific computations in distributed computing environments. We explore an approach based on separation of computation design (application development) and distributed execution of computations, and investigate best practices for construction of virtual infrastructures for computational science - software systems that abstract and virtualize the processes of managing scientific computations on heterogeneous distributed resource systems. As a result we present StratUm, a toolkit for management of eScience computations. To illustrate use of the toolkit, we present it in the context of a case study where we extend the capabilities of an existing kinetic Monte Carlo software framework to utilize distributed computational resources. The case study illustrates a viable design pattern for construction of virtual infrastructures for distributed scientific computing. The resulting infrastructure is evaluated using a computational experiment from molecular systems biology. Per-Olov Östberg, Andreas Hellander, Brian Drawert, Erik Elmroth, Sverker Holmgren, Linda R. Petzold |
CCGRID | 2 |
| 2010 | CellMC - a multiplatform model compiler for the Cell Broadband Engine and x86abstractMOTIVATION: Gillespie's stochastic simulation algorithm (SSA) is often the most tractable method to study stochastic models of biochemical systems. The algorithm itself is very simple and a natural target for implementation on specialized architectures such as the Cell Broadband Engine (Cell/BE). We have developed CellMC, a multiplatform SBML model compiler implementing a vectorized version of SSA for use on Cell/BE or x86 PCs. AVAILABILITY: The code is freely available from http://www.cellmc.org. It will run on a wide variety of x86 computers running Linux/MacOSX (Darwin) and on Cell/BE computers such as the Sony PlayStation3 (PS3) and the IBM BladeCenter QS22. CellMC requires gcc, libxml2 and libxslt, all of which are installed by default on most of the supported platforms. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Emmet Caulfield, Andreas Hellander |
Bioinform. | 2 |