Daniel Gillblad

dblp:48/5973 · DBLP profile ↗
← Back
22ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0001-8952-3542ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 1 since 2021Computer networks · 6Databases, data management, data science and information retrieval · 6 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2023 Exploring Trade-Offs in MLOps Adoption
abstract
Machine Learning Operations (MLOps) play a crucial role in the success of data science projects in companies. However, despite its obvious benefits, several companies struggle to adopt MLOps practices and face difficulty in deciding how to deploy and evolve ML models. To gain a deeper understanding of these challenges, we conduct a multi-case study involving nine practitioners from seven companies. Based on our empirical results, we identify the key trade-offs we see companies make when adopting MLOps. We categorise these trade-offs into four concerns of the BAPO model: Business, Architecture, Process, and Organisation. Finally, we provide suggestions to mitigate the identified trade-offs. By identifying and detailing these trade-offs and the implications of these, this research helps companies to ensure the successful adoption of MLOps.
Meenu Mary John, Helena Olsson, Jan Bosch, Daniel Gillblad
APSEC4
2023 Advancing MLOps from Ad hoc to Kaizen
abstract
Companies across various domains increasingly adopt Machine Learning Operations (MLOps) as they recognise the significance of operationalising ML models. Despite growing interest from practitioners and ongoing research, MLOps adoption in practice is still in its initial stages. To explore the adoption of MLOps, we employ a multi-case study in seven companies. Based on empirical findings, we propose a maturity model outlining the typical stages companies undergo when adopting MLOps, ranging from Ad hoc to Kaizen. We identify five dimensions associated with each stage of the maturity model as part of our MLOps framework. We also map these seven companies to the identified stages in the maturity model. Our study serves as a roadmap for companies to assess their current state of MLOps, identify gaps and overcome obstacles to successfully adopting MLOps.
Meenu Mary John, Daniel Gillblad, Helena Olsson, Jan Bosch
SEAA2
2021 Adversarial representation learning for synthetic replacement of private attributes
abstract
Data privacy is an increasingly important aspect of many real-world analytics tasks. Data sources that contain sensitive information may have immense potential which could be unlocked using the right privacy enhancing transformations, but current methods often fail to produce convincing output. Furthermore, finding the right balance between privacy and utility is often a tricky trade-off. In this work, we propose a novel approach for data privatization, which involves two steps: in the first step, it removes the sensitive information, and in the second step, it replaces this information with an independent random sample. Our method builds on adversarial representation learning which ensures strong privacy by training the model to fool an increasingly strong adversary. While previous methods only aim at obfuscating the sensitive information, we find that adding new random information in its place strengthens the provided privacy and provides better utility at any given level of privacy. The result is an approach that can provide stronger privatization on image data, and yet be preserving both the domain and the utility of the inputs, entirely independent of the downstream task.
John Martinsson, Edvin Listo Zec, Daniel Gillblad, Olof Mogren
IEEE BigData3
2018 Streaming word similarity mining on the cheap
abstract
Accurately and efficiently estimating word similarities from text is fundamental in natural language processing.In this paper, we propose a fast and lightweight method for estimating similarities from streams by explicitly counting second-order co-occurrences.The method rests on the observation that words that are highly correlated with respect to such counts are also highly similar with respect to firstorder co-occurrences.Using buffers of cooccurred words per word to count secondorder co-occurrences, we can then estimate similarities in a single pass over data without having to do prohibitively expensive similarity calculations.We demonstrate that this approach is scalable, converges rapidly, behaves robustly under parameter changes, and that it captures word similarities on par with those given by state-of-the-art word embeddings.
Olof Görnerup, Daniel Gillblad
EMNLP2
2017 Domain-agnostic discovery of similarities and concepts at scale
Olof Görnerup, Daniel Gillblad, Theodore Vasiloudis
Knowl. Inf. Syst.2
2015 Predicting Swedish Elections with Twitter: A Case for Stochastic Link Structure Analysis
abstract
The question that whether Twitter data can be leveraged to forecast outcome of the elections has always been of great anticipation in the research community. Existing research focuses on leveraging content analysis for positivity or negativity analysis of the sentiments of opinions expressed. This is while, analysis of link structure features of social networks underlying the conversation involving politicians has been less looked. The intuition behind such study comes from the fact that density of conversations about parties along with their respective members, whether explicit or implicit, should reflect on their popularity. On the other hand, dynamism of interactions, can capture the inherent shift in popularity of accounts of politicians. Within this manuscript we present evidence of how a well-known link prediction algorithm, can reveal an authoritative structural link formation within which the popularity of the political accounts along with their neighbourhoods, shows strong correlation with the standing of electoral outcomes. As an evidence, the public time-lines of two electoral events from 2014 elections of Sweden on Twitter have been studied. By distinguishing between member and official party accounts, we report that even using a focus-crawled public dataset, structural link popularities bear strong statistical similarities with vote outcomes. In addition we report strong ranked dependence between standings of selected politicians and general election outcome, as well as for official party accounts and European election outcome.
Nima Dokoohaki, Filippia Zikou, Daniel Gillblad, Mihhail Matskin
ASONAM3
2015 Predicting service metrics for cluster-based services using real-time analytics
abstract
Predicting the performance of cloud services is intrinsically hard. In this work, we pursue an approach based upon statistical learning, whereby the behaviour of a system is learned from observations. Specifically, our testbed implementation collects device statistics from a server cluster and uses a regression method that accurately predicts, in real-time, client-side service metrics for a video streaming service running on the cluster. The method is service-agnostic in the sense that it takes as input operating-systems statistics instead of service-level metrics. We show that feature set reduction significantly improves prediction accuracy in our case, while simultaneously reducing model computation time. We also discuss design and implementation of a real-time analytics engine, which processes streams of device statistics and service metrics from testbed sensors and produces model predictions through online learning.
Rerngvit Yanggratoke, Jawwad Ahmed, John Ardelius, Christofer Flinta, Andreas Johnsson, Daniel Gillblad, Rolf Stadler
CNSM6
2015 Exploring communication and mobility behavior of 3G network users and its temporal consistency
abstract
Over the past decade, telecommunication network operators have more and more realized the added value of data analytics for their network deployment efficiency. Early studies targeted the global network perspective by localizing peak loads, both in terms of area and time period. Due to their higher granularity and information richness, current telecommunication datasets allow increasingly deeper insights into the network activities of the users. Existing network traffic classification studies tend to divide users into groups without considering the transitions between different groups caused by individual behavioral traits, which we expect to show observable regularities. Our approach defines a profiling model that characterizes the user behavior as well as its temporal dynamics from two perspectives: w.r.t. (i) the network load the users generate, and (ii) their mobility patterns. The model is evaluated with two unsupervised clustering algorithms of different complexity (namely, XMeans and EM) by means of a 3G trace dataset from a European operator.
Andrea Hess, Ian Marsh, Daniel Gillblad
ICC3
2015 Knowing an Object by the Company it Keeps: A Domain-Agnostic Scheme for Similarity Discovery
abstract
Appropriately defining and then efficiently calculating similarities from large data sets are often essential in data mining, both for building tractable representations and for gaining understanding of data and generating processes. Here we rely on the premise that given a set of objects and their correlations, each object is characterized by its context, i.e. its correlations to the other objects, and that the similarity between two objects therefore can be expressed in terms of the similarity between their respective contexts. Resting on this principle, we propose a data-driven and highly scalable approach for discovering similarities from large data sets by representing objects and their relations as a correlation graph that is transformed to a similarity graph. Together these graphs can express rich structural properties among objects. Specifically, we show that concepts -- representations of abstract ideas and notions -- are constituted by groups of similar objects that can be identified by clustering the objects in the similarity graph. These principles and methods are applicable in a wide range of domains, and will here be demonstrated for three distinct types of objects: codons, artists and words, where the numbers of objects and correlations range from small to very large.
Olof Görnerup, Daniel Gillblad, Theodore Vasiloudis
ICDM2
2015 Predicting real-time service-level metrics from device statistics
abstract
While real-time service assurance is critical for emerging telecom cloud services, understanding and predicting performance metrics for such services is hard. In this paper, we pursue an approach based upon statistical learning whereby the behavior of the target system is learned from observations. We use methods that learn from device statistics and predict metrics for services running on these devices. Specifically, we collect statistics from a Linux kernel of a server machine and predict client-side metrics for a video-streaming service (VLC). The fact that we collect thousands of kernel variables, while omitting service instrumentation, makes our approach service-independent and unique. While our current lab configuration is simple, our results, gained through extensive experimentation, prove the feasibility of accurately predicting client-side metrics, such as video frame rates and RTP packet rates, often within 10-15% error (NMAE), also under high computational load and across traces from different scenarios.
Rerngvit Yanggratoke, Jawwad Ahmed, John Ardelius, Christofer Flinta, Andreas Johnsson, Daniel Gillblad, Rolf Stadler
IM6
2015 A platform for predicting real-time service-level metrics from device statistics
abstract
Predicting performance metrics for cloud services is critical for real-time service assurance. We demonstrate a platform for estimating real-time service-level metrics. Statistical learning methods on device statistics are used to predict metrics for services running on these devices.
Rerngvit Yanggratoke, Jawwad Ahmed, John Ardelius, Christofer Flinta, Andreas Johnsson, Daniel Gillblad, Rolf Stadler
IM6
2015 Autonomous Load Balancing of Heterogeneous Networks
abstract
This paper presents a method for load balancing heterogeneous networks by dynamically assigning values to the LTE cell range expansion (CRE) parameter. The method records hand-over events online and adapts flexibly to changes in terminal traffic and mobility by maintaining statistical estimators that are used to support autonomous assignment decisions. The proposed approach has low overhead and is highly scalable due to a modularised and completely distributed design that exploits self-organisation based on local inter-cell interactions. An advanced simulator that incorporates terminal traffic patterns and mobility models with a radio access network simulator has been developed to validate and evaluate the method.
Per Kreuger, Olof Görnerup, Daniel Gillblad, Tomas Lundborg, Diarmuid Corcoran, Andreas Ermedahl
VTC Spring3
2014 Learning machines for computational epidemiology
abstract
Resting on our experience of computational epidemiology in practice and of industrial projects on analytics of complex networks, we point to an innovation opportunity for improving the digital services to epidemiologists for monitoring, modeling, and mitigating the effects of communicable disease. Artificial intelligence and intelligent analytics of syndromic surveillance data promise new insights to epidemiologists, but the real value can only be realized if human assessments are paired with assessments made by machines. Neither massive data itself, nor careful analytics will necessarily lead to better informed decisions. The process producing feedback to humans on decision making informed by machines can be reversed to consider feedback to machines on decision making informed by humans, enabling learning machines. We predict and argue for the fact that the sensemaking that such machines can perform in tandem with humans can be of immense value to epidemiologists in the future.
Magnus Boman, Daniel Gillblad
IEEE BigData2
2014 Explaining Probabilistic Fault Diagnosis and Classification Using Case-Based Reasoning
Tomas Olsson, Daniel Gillblad, Peter Funk 0001, Ning Xiong 0001
ICCBR2
2012 Performance evaluation of a distributed and probabilistic network monitoring approach
Rebecca Steinert, Daniel Gillblad
CNSM2
2012 Zero Configuration Adaptive Paging (zCap)
abstract
Today, cellular networks rely on fixed collections of cells (tracking areas) for handset localisation. This management parameter is manually configured and maintained and is not regularly adapted to changes in use patterns. We present a decentralised approach to localisation, based on a self-adaptive probabilistic mobility model. Estimates of model parameters are built from observations of mobility patterns collected on-line using a distributed algorithm. Based on these estimates, dynamic local neighbourhoods of cells are formed and maintained by the mobility management entities of the network. These neighbourhoods replace the static tracking areas used in current implementations by using the tracking area list facility of LTE. The model is also used to derive a multi phase paging scheme, where the division of cells into consecutive phases is optimal with respect to a set balance between response times and paging cost. The approach requires no manual tracking area configuration, and performs localisation efficiently in terms of number of location updates, page messages per localisation request and response times.
Per Kreuger, Daniel Gillblad, Åke Arvidsson
VTC Fall2
2011 Translation of probabilistic QoS in hierarchical and decentralized settings
abstract
In this paper we build on methods from probabilistic management for overcoming two issues in the translation of QoS into configurations of network nodes in dynamic, decentralized, and hierarchical networks. First, the inherent uncertainty about node performance in such networks (due to network dynamics) may impede adequate specification of QoS. We suggest how a probabilistic variant of QoS via optimization can be translated into objectives that can be handled with probabilistic management methods. Second, the dynamics of the network may make the interpretation of QoS dependent on the availability of network resources. We use methods from probabilistic management to suggest how the expression of QoS may be restricted in order to be able to ensure the existence of a translation of given QoS. The suggestions are applied to our own network architecture. The aim is to further develop our approach to a viable management solution for cloud computing and virtualized networks.
Björn Bjurling, Rebecca Steinert, Daniel Gillblad
APNOMS3
2011 A Distributed Spatio-Temporal Event Correlation Protocol for Multi-Layer Virtual Networks
abstract
We present a distributed spatio-temporal event correlation protocol for multi-layer networks. The problems that we address relate to scalability in stacked overlay networks and network equipment with asynchronous clocks, which complicates the problem of event correlation. We describe a cross-layer protocol designed to address these problems, operating in a fully distributed manner and taking into account asynchronous timestamps. It is assumed that events in one layer may arise from a series of events in lower layers. Detected events that are spatially related in one layer are aggregated using a gossip-like protocol, and constitute a root cause. The set of aggregated events is disseminated to lower layers and used for temporal correlation. We have tested the scalability and the performance of the distributed event protocol, using both synthetically generated and real-world topologies. The results indicate that the average overhead produced for collecting events down the stack of overlays increases with the number of layers. For a fixed number of layers, the protocol scales similarly with the graph-theoretic properties for a network of increasing size.
Rebecca Steinert, S. Gestrelius, Daniel Gillblad
GLOBECOM3
2010 Long-Term Adaptation and Distributed Detection of Local Network Changes
abstract
We present a statistical approach to distributed detection of local latency shifts in networked systems. For this purpose, response delay measurements are performed between neighbouring nodes via probing. The expected probe response delay on each connection is statistically modelled via parameter estimation. Adaptation to drifting delays is accounted for by the use of overlapping models, such that previous models are partially used as input to future models. Based on the symmetric Kullback-Leibler divergence metric, latency shifts can be detected by comparing the estimated parameters of the current and previous models. In order to reduce the number of detection alarms, thresholds for divergence and convergence are used. The method that we propose can be applied to many types of statistical distributions, and requires only constant memory compared to e.g., sliding window techniques and decay functions. Therefore, the method is applicable in various kinds of network equipment with limited capacity, such as sensor networks, mobile ad hoc networks etc. We have investigated the behaviour of the method for different model parameters. Further, we have tested the detection performance in network simulations, for both gradual and abrupt shifts in the probe response delay. The results indicate that over 90% of the shifts can be detected. Undetected shifts are mainly the effects of long convergence processes triggered by previous shifts. The overall performance depends on the characteristics of the shifts and the configuration of the model parameters.
Rebecca Steinert, Daniel Gillblad
GLOBECOM2
2009 Discovering Process Models from Unlabelled Event Logs
Diogo R. Ferreira, Daniel Gillblad
BPM2
2005 Emulating Process Simulators with Learning Systems
Daniel Gillblad, Anders Holst, Björn Levin
ICANN (2)1
2001 Dependency Derivation in Industrial Process Data
abstract
In many industrial processes, finding dependencies and the creation of dependency graphs can increase the understanding of the system significantly. This knowledge can then be used for further optimization and variable selection. Most of the measured attributes in these cases come in the form of time series. There are several ways of determining correlation between series, most of them suffering from specific problems when applied to real-world data. Here, a well performing measure based on the mutual information rate is derived and discussed with results from both synthetic and real data.
Daniel Gillblad, Anders Holst
ICDM1