VLDB 2026 Research / reviewers in the wild / expert
Salman Zubair Toor
dblp:13/8188 · also Salman Toor
· DBLP profile ↗
20ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0003-0302-6276ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 10 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 5 · 4 since 2021Artificial intelligence and machine learning · 4 · 3 since 2021Software engineering, systems software and programming languages · 4 · 2 first-authorSystems, architecture and hardware · 3 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Quantifying Catastrophic Forgetting in IoT Intrusion Detection Systems
Sourasekhar Banerjee, David Bergqvist, Salman Zubair Toor, Christian Rohner, Andreas Johnsson |
ICC | 3 |
| 2024 | GNN-IDS: Graph Neural Network based Intrusion Detection SystemabstractIntrusion detection systems (IDSs) are widely used to identify anomalies in computer networks and raise alarms on intrusive behaviors. ML-based IDSs generally take network traces or host logs as input to extract patterns from individual samples, whereas the inter-dependencies of network are often not captured and learned, which may result in large amounts of uncertain predictions, false positives, and false negatives. To tackle the challenges in intrusion detection, we propose a graph neural network-based intrusion detection system (GNN-IDS), which is data-driven and machine learning-empowered. In our proposed GNN-IDS, the attack graph and real-time measurements that represent static and dynamic attributes of computer networks, respectively, are incorporated and associated to represent complex computer networks. Graph neural networks are employed as the inference engine for intrusion detection. By learning network connectivity, graph neural networks can quantify the importance of neighboring nodes and node features to make more reliable predictions. Furthermore, by incorporating an attack graph, GNN-IDS could not only detect anomalies but also identify the malicious actions causing the anomalies. The experimental results on a use case network with two synthetic datasets (one generated from public IDS data) show that the proposed GNN-IDS achieves good performance. The results are analyzed from the aspects of uncertainty, explainability, and robustness. Zhenlu Sun, André Teixeira 0001, Salman Zubair Toor |
ARES | 3 |
| 2024 | Empowering Data Mesh with Federated LearningabstractThe evolution of data architecture has seen the rise of data lakes, aiming to solve the bottlenecks of data management and promote intelligent decision-making. However, this centralized architecture is limited by the proliferation of data sources and the growing demand for timely analysis and processing. A new data paradigm, Data Mesh, is proposed to overcome these challenges. In this decentralized architecture where data is locally preserved by each domain team, traditional centralized machine learning cannot conduct effective analysis across multiple domains, especially for security-sensitive organizations. To this end, we introduce a pioneering approach that incorporates Federated Learning into Data Mesh. This applied research article emphasizes the benefits of combining two distinct domains to achieve the best outcomes for industrial use cases. Salman Zubair Toor |
IEEE Big Data | 2 |
| 2024 | Data management of scientific applications in a reinforcement learning-based hierarchical storage systemabstractIn many areas of data-driven science, large datasets are generated where the individual data objects are images, matrices, or otherwise have a clear structure. However, these objects can be information-sparse, and a challenge is to efficiently find and work with the most interesting data as early as possible in an analysis pipeline. We have recently proposed a new model for big data management where the internal structure and information of the data are associated with each data object (as opposed to simple metadata). There is then an opportunity for comprehensive data management solutions to account for data-specific internal structure as well as access patterns. In this article, we explore this idea together with our recently proposed hierarchical storage management framework that uses reinforcement learning (RL) for autonomous and dynamic data placement in different tiers in a storage hierarchy. Our case-study is based on four scientific datasets: Protein translocation microscopy images, Airfoil angle of attack meshes, 1000 Genomes sequences, and Phenotypic screening images. The presented results highlight that our framework is optimal and can quickly adapt to new data access requirements. It overall reduces the data processing time, and the proposed autonomous data placement is superior compared to any static or semi-static data placement policies. Tianru Zhang, Ankit Gupta 0018, María Andreína Francisco Rodríguez, Ola Spjuth, Andreas Hellander, Salman Zubair Toor |
Expert Syst. Appl. | 6 |
| 2023 | Efficient Resource Scheduling for Distributed Infrastructures Using Negotiation CapabilitiesabstractThe information explosion drives enterprises and individuals to rent cloud computing infrastructure for their applications in the cloud. However, the agreements between cloud computing providers and clients are often inefficient. We propose an agent-based auto-negotiation system for resource scheduling using fuzzy logic. Our method completes a one-to-one auto-negotiation process and generates optimal offers for providers and clients. We compare the impact of different member functions, fuzzy rule sets, and negotiation scenarios on the offers to optimize the system. Our proposed method efficiently utilizes resources and offers interpretability, high flexibility, and customization. We successfully train machine learning models to replace the fuzzy negotiation system, improving processing speed. The article also highlights potential future improvements to the proposed system and machine learning models. All codes and data are available as an open source repository. Junjie Chu 0002, Salman Zubair Toor |
CLOUD | 3 |
| 2023 | Efficient Hierarchical Storage Management Empowered by Reinforcement Learning Extended AbstractabstractWith the rapid development of big data and cloud computing, data management has become increasingly challenging. A possible solution is to use an intelligent hierarchical (multi-tier) storage system (HSS). An HSS is a meta solution that consists of different storage frameworks organized as a jointly constructed storage pool. A built-in data migration policy that determines the optimal placement of the datasets in the hierarchy is essential. Placement decisions are a non-trivial task since they should be made according to the characteristics of the dataset, the tier status in a hierarchy, and access patterns. This paper presents an open-source hierarchical storage framework with a dynamic migration policy based on reinforcement learning (RL). Tianru Zhang, Andreas Hellander, Salman Zubair Toor |
ICDE | 3 |
| 2023 | Efficient Hierarchical Storage Management Empowered by Reinforcement LearningabstractWith the rapid development of big data and cloud computing, data management has become increasingly challenging. Over the years, a number of frameworks for data management have become available. Most of them are highly efficient, but ultimately create data silos. It becomes difficult to move and work coherently with data as new requirements emerge. A possible solution is to use an intelligent hierarchical (multi-tier) storage system (HSS). A HSS is a meta solution that consists of different storage frameworks organized as a jointly constructed storage pool. A built-in data migration policy that determines the optimal placement of the datasets in the hierarchy is essential. Placement decisions is a non-trivial task since it should be made according to the characteristics of the dataset, the tier status in a hierarchy, and access patterns. This paper presents an open-source hierarchical storage framework with a dynamic migration policy based on reinforcement learning (RL). We present a mathematical model, a software architecture, and implementations based on both simulations and a live cloud-based environment. We compare the proposed RL-based strategy to a baseline of three rule-based policies, showing that the RL-based policy achieves significantly higher efficiency and optimal data distribution in different scenarios. Tianru Zhang, Andreas Hellander, Salman Zubair Toor |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | To test, or not to test: A proactive approach for deciding complete performance test initiationabstractSoftware performance testing requires a set of inputs that exercise different sections of the code to identify performance issues. However, running tests on a large set of inputs can be a very time consuming process. It is even more problematic when test inputs are constantly growing, which is the case with a large-scale scientific organization such as CERN where the process of performing scientific experiment generates plethora of data that is analyzed by physicists leading to new scientific discoveries. Therefore, in this article, we present a test input minimization approach based on a clustering technique to handle the issue of testing on growing data. Furthermore, we use clustering information to propose an automatic approach that recommends the tester to decide when to run the complete test suite for performance testing. To demonstrate the efficacy of our approach, we applied it to two different code updates of a web service which is used at CERN and we found that the recommendation for performance test initiation made by our approach for an update with bottleneck is valid. Omar Javed, Giles Reger, Salman Zubair Toor |
IEEE Big Data | 4 |
| 2022 | Scalable federated machine learning with FEDnabstractFederated machine learning promises to overcome the input privacy challenge in machine learning. By iteratively updating a model on private clients and aggregating these local model updates into a global federated model, private data is incorporated in the federated model without needing to share and expose that data. Several open software projects for federated learning have appeared. Most of them focuses on supporting flexible experimentation with different model aggregation schemes and with different privacy-enhancing technologies. However, there is a lack of open frameworks that focuses on critical distributed computing aspects of the problem such as scalability and resilience. It is a big step to take for a data scientist to go from an experimental sandbox to testing their federated schemes at scale in real-world geographically distributed settings. To bridge this gap we have designed and developed a production-grade hierarchical federated learning framework, FEDn. The framework is specifically designed to make it easy to go from local development in pseudo-distributed mode to horizontally scalable distributed deployments. FEDn both aims to be production grade for industrial applications and a flexible research tool to explore real-world performance of novel federated algorithms and the framework has been used in number of industrial and academic R&D projects. In this paper we present the architecture and implementation of FEDn. We demonstrate the framework's scalability and efficiency in evaluations based on two case-studies representative for a cross-silo and a cross-device use-case respectively. Morgan Ekmefjord, Addi Ait-Mlouk, Sadi Alawadi, Mattias Åkesson, Ola Spjuth, Salman Zubair Toor, Andreas Hellander |
CCGRID | 7 |
| 2020 | Smart Resource Management for Data Streaming using an Online Bin-packing StrategyabstractData stream processing frameworks provide reliable and efficient mechanisms for executing complex workflows over large datasets. A common challenge for the majority of currently available streaming frameworks is efficient utilization of resources. Most frameworks use static or semi-static settings for resource utilization that work well for established use cases but lead to marginal improvements for unseen scenarios. Another pressing issue is the efficient processing of large individual objects such as images and matrices typical for scientific datasets. HarmonicIO has proven to be a good solution for streams of relatively large individual objects, as demonstrated in a benchmark comparison with the Apache Spark and Kafka streaming frameworks. We here present an extension of the HarmonicIO framework based on the online bin-packing algorithm. The main focus is to compare different strategies adapted in streaming frameworks for efficient resource utilization. Based on a real world use case from large-scale microscopy pipelines, we compare two different strategies of auto-scaling implemented in the HarmonicIO and Spark Streaming frameworks. Oliver Stein, Ben Blamey, Alan Sabirsh, Ola Spjuth, Andreas Hellander, Salman Zubair Toor |
IEEE BigData | 7 |
| 2019 | Adapting the Secretary Hiring Problem for Optimal Hot-Cold Tier Placement Under Top-K WorkloadsabstractTop-K queries are an established heuristic in information retrieval. This paper presents an approach for optimal tiered storage allocation under stream processing workloads using this heuristic: those requiring the analysis of only the top-K ranked most relevant documents from a fixed-length stream, stream window, or batch job. Documents are ranked for relevance on a user-specified interestingness function, the top-K stored for further processing. This scenario bears similarity to the classic Secretary Hiring Problem (SHP), and the expected rate of document writes and document lifetime can be modelled as a function of document index. We present parameter-based algorithms for storage tier placement, minimizing document storage and transport costs. We derive expressions for optimal parameter values in terms of tier storage and transport costs a priori, without needing to monitor the application. This contrasts with (often complex) existing work on tiered storage optimization, which is either tightly coupled to specific use cases, or requires active monitoring of application IO load - ill-suited to long-running or one-off operations common in the scientific computing domain. We motivate and evaluate our model with a trace-driven simulation of human-in-the-loop bio-chemical model exploration, and two cloud storage case studies. Ben Blamey, Fredrik Wrede, Andreas Hellander, Salman Zubair Toor |
CCGRID | 5 |
| 2018 | HarmonicIO: Scalable Data Stream Processing for Scientific DatasetsabstractMany streaming frameworks have been introduced to deal with the needs for online analysis of massive datasets. Scientific applications often require significant changes to make them compatible with these frameworks. Other issues include tight coupling with the underlying infrastructure, shared computing environment, static topology settings, and complex configuration. In this article we present HarmonicIO, a lightweight streaming framework specialized for scientific datasets. It boasts a smart dynamic architecture, is highly elastic, and enforces a clear separation between framework components and application execution environment using container technology. Preechakorn Torruangwatthana, Håkan Wieslander, Ben Blamey, Andreas Hellander, Salman Zubair Toor |
IEEE CLOUD | 5 |
| 2018 | BAMSI: a multi-cloud service for scalable distributed filtering of massive genome dataabstractBACKGROUND: The advent of next-generation sequencing (NGS) has made whole-genome sequencing of cohorts of individuals a reality. Primary datasets of raw or aligned reads of this sort can get very large. For scientific questions where curated called variants are not sufficient, the sheer size of the datasets makes analysis prohibitively expensive. In order to make re-analysis of such data feasible without the need to have access to a large-scale computing facility, we have developed a highly scalable, storage-agnostic framework, an associated API and an easy-to-use web user interface to execute custom filters on large genomic datasets. RESULTS: We present BAMSI, a Software as-a Service (SaaS) solution for filtering of the 1000 Genomes phase 3 set of aligned reads, with the possibility of extension and customization to other sets of files. Unique to our solution is the capability of simultaneously utilizing many different mirrors of the data to increase the speed of the analysis. In particular, if the data is available in private or public clouds - an increasingly common scenario for both academic and commercial cloud providers - our framework allows for seamless deployment of filtering workers close to data. We show results indicating that such a setup improves the horizontal scalability of the system, and present a possible use case of the framework by performing an analysis of structural variation in the 1000 Genomes data set. CONCLUSIONS: BAMSI constitutes a framework for efficient filtering of large genomic data sets that is flexible in the use of compute as well as storage resources. The data resulting from the filter is assumed to be greatly reduced in size, and can easily be downloaded or routed into e.g. a Hadoop cluster for subsequent interactive analysis using Hive, Spark or similar tools. In this respect, our framework also suggests a general model for making very large datasets of high scientific value more accessible by offering the possibility for organizations to share the cost of hosting data on hot storage, without compromising the scalability of downstream analysis. Kristiina Ausmees, Aji John, Salman Zubair Toor, Andreas Hellander, Carl Nettelblad |
BMC Bioinform. | 3 |
| 2018 | Secure Cloud Connectivity for Scientific ApplicationsabstractCloud computing improves utilization and flexibility in allocating computing resources while reducing the infrastructural costs. However, in many cases cloud technology is still proprietary and tainted by security issues rooted in the multi-user and hybrid cloud environment. A lack of secure connectivity in a hybrid cloud environment hinders the adaptation of clouds by scientific communities that require scaling-out of the local infrastructure using publicly available resources for large-scale experiments. In this article, we present a case study of the DII-HEP secure cloud infrastructure and propose an approach to securely scale-out a private cloud deployment to public clouds in order to support hybrid cloud scenarios. A challenge in such scenarios is that cloud vendors may offer varying and possibly incompatible ways to isolate and interconnect virtual machines located in different cloud networks. Our approach is tenant driven in the sense that the tenant provides its connectivity mechanism. We provide a qualitative and quantitative analysis of a number of alternatives to solve this problem. We have chosen one of the standardized alternatives, Host Identity Protocol, for further experimentation in a production system because it supports legacy applications in a topologically-independent and secure way. Lirim Osmani, Salman Zubair Toor, Miika Komu, Matti J. Kortelainen, Tomas Lindén, Rasib Hassan Khan, Paula Eerola, Sasu Tarkoma |
IEEE Trans. Serv. Comput. | 2 |
| 2017 | Cost-aware Application Development and Management using CLOUD-METRIC
Alieu Jallow, Andreas Hellander, Salman Zubair Toor |
CLOSER | 3 |
| 2017 | SNIC Science Cloud (SSC): A National-Scale Cloud Infrastructure for Swedish AcademiaabstractThe cloud computing paradigm have fundamentally changed the way computational resources are being offered. Although the number of large-scale providers in academia is still relatively small, there is a rapidly increasing interest and adoption of cloud Infrastructure-as-a-Service in the scientific community. The added flexibility in how applications can be implemented compared to traditional batch computing systems is one of the key success factors for the paradigm, and scientific cloud computing promises to increase adoption of simulation and data analysis in scientific communities not traditionally users of large scale e-Infrastructure, the so called ”long tail of science”. In 2014, the Swedish National Infrastructure for Computing (SNIC) initiated a project to investigate the cost and constraints of offering cloud infrastructure for Swedish academia. The aim was to build a platform where academics could evaluate cloud computing for their use-cases. SNIC Science Cloud (SSC) has since then evolved into a national-scale cloud infrastructure based on three geographically distributed regions. In this article we present the SSC vision, architectural details and user stories. We summarize the experiences gained from running a nationalscale cloud facility into ”ten simple rules” for starting up a science cloud project based on OpenStack. We also highlight some key areas that require careful attention in order to offer cloud infrastructure for ubiquitous academic needs and in particular scientific workloads. Salman Zubair Toor, Mathias Lindberg, Ingemar Falman, Andreas Vallin, Olof Mohill, Pontus Freyhult, Linus Nilsson, Martin Agback, Lars Viklund, Henric Zazzik, Ola Spjuth, Marco Capuccini, Joakim Moller, Donal Murtagh, Andreas Hellander |
eScience | 1 |
| 2017 | A Flexible Computational Framework Using R and Map-Reduce for Permutation Tests of Massive Genetic Analysis of Complex TraitsabstractIn quantitative trait locus (QTL) mapping significance of putative QTL is often determined using permutation testing. The computational needs to calculate the significance level are immense, 104up to 108or even more permutations can be needed. We have previously introduced the PruneDIRECT algorithm for multiple QTL scan with epistatic interactions. This algorithm has specific strengths for permutation testing. Here, we present a flexible, parallel computing framework for identifying multiple interacting QTL using the PruneDIRECT algorithm which uses the map-reduce model as implemented in Hadoop. The framework is implemented in R, a widely used software tool among geneticists. This enables users to rearrange algorithmic steps to adapt genetic models, search algorithms, and parallelization steps to their needs in a flexible way. Our work underlines the maturity of accessing distributed parallel computing for computationally demanding bioinformatics applications through building workflows within existing scientific environments. We investigate the PruneDIRECT algorithm, comparing its performance to exhaustive search and DIRECT algorithm using our framework on a public cloud resource. We find that PruneDIRECT is vastly superiorfor permutation testing, and perform 2 × 105permutations for a 2D QTL problem in 15 hours, using 100 cloud processes. We show that our framework scales out almost linearly for a 3D QTL search. Behrang Mahjani, Salman Zubair Toor, Carl Nettelblad, Sverker Holmgren |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | Scientific Analysis by Queries in Extended SPARQL over a Scalable e-Science Data StoreabstractData-intensive applications in e-Science require scalable solutions for storage as well as interactive tools for analysis of scientific data. It is important to be able to query the data in a storage-independent way, and to be able to obtain the results of the data-analysis incrementally (in contrast to traditional batch solutions). We use the RDF data model extended with multidimensional numeric arrays to represent the results, parameters, and other metadata describing scientific experiments, and SciSPARQL, an extension of the SPARQL language, to combine massive numeric array data and metadata in queries. To address the scalability problem we present an architecture that enables the same SciSPARQL queries to be executed on the RDF dataset whether it is stored in a relational DBMS or mapped over a specialized geographically distributed e-Science data store. In order to minimize access and communication costs, we represent the arrays with proxy objects, and retrieve their content lazily. We formulate typical analysis tasks from a computational biology application in terms of SciSPARQL queries, and compare the query processing performance with manually written scripts in MATLAB. Andrej Andrejev, Salman Zubair Toor, Andreas Hellander, Sverker Holmgren, Tore Risch |
e-Science | 2 |
| 2012 | Investigating an Open Source Cloud Storage Infrastructure for CERN-specific Data AnalysisabstractWe present a first case study where an open source storage cloud based on Openstack - SWIFT is used for handling data from CERN experiments using the ROOT software framework. This type of storage clouds promise to be easy to deploy and provide transparent access to data using standardized protocols. We examine the scalability and performance of the system using test cases which are derived from the normal usage and the structure of the ROOT software. The results show that cloud solutions like the SWIFT storage system could fulfill the requirements by the CERN scientific community. To verify this, a more extensive effort with many more tests and use-cases is needed. However, the impact of providing alternate storage solutions is large and further work is motivated. Salman Zubair Toor, Rainer Töebbicke, Maitane Zotes Resines, Sverker Holmgren |
NAS | 1 |
| 2011 | A Scalable Architecture for e-Science Data ManagementabstractThe massive increase in the size of the data provided by e-Science applications requires not only to increase the capabilities of resources, but also to design new strategies for efficient utilization of already available resources. In this paper we present a scalable approach to extend a file-oriented storage system, Chelonia, with geographically distributed databases defined by a generic database schema. The database schema is able to model the data from typical e-Science applications. The system includes web service query service allowing e-Science applications to query the required data. Salman Zubair Toor, Manivasakan Sabesan, Sverker Holmgren, Tore Risch |
eScience | 1 |