David Hästbacka

dblp:18/1537 · DBLP profile ↗
← Back
26ranked-venue papers
9as first author
16since 2021 · last 2026
0000-0001-8442-1248ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 6 first-author · 6 since 2021Software engineering, systems software and programming languages · 10 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author
YearPublicationVenuePosition
2026 Designing a cloud-native MLOps pipeline using Databricks: A case study in industrial practice
abstract
We present a engineering case study on the design and evolution of a cloud-native MLOps pipeline in an industrial setting using Databricks. Motivated by the operational and lifecycle challenges of deploying machine-learning models at scale, the project adopted a platform-centric approach to automation, reproducibility, and scalable operation. The pipeline integrates versioned data management, experiment tracking and registry-based promotion, CI/CD for ML, and environment isolation to bridge experimentation and production. Security and data-privacy constraints are operationalized through workspace isolation, role-based access control, and secrets management integrated into deployment automation. Our methodological approach entailed an engineering case-study design, incorporating triangulation across development artifacts, CI/CD records, and collaborative design episodes. The paper delivers three contributions: (i) a platform-grounded reference architecture documenting the design decisions, trade-offs, and deliberate deviations from vendor guidance that shaped the final implementation; (ii) collaboration practices that align roles across DataOps, ModelOps, and DevOps showing how shared artifacts and platform constraints structure cross-role coordination; and (iii) recurring implementation patterns and practitioners lessons that are capability-oriented and transferable beyond the specific platform. The results provide practitioners with empirical evidence on how MLOps automation is realized and constrained in practice, filling a gap that vendor documentation and conceptual frameworks alone do not address.
Sergio Moreschini, Sandra Raitaniemi, Elias Mäkelä, Tommi Laukkanen, David Hästbacka
Future Gener. Comput. Syst.5
2026 Computation offloading in collaborative computing systems: A comprehensive survey on methods, challenges, and future directions
abstract
• Complex offloading dynamics are rarely considered in offloading problems • The potential of collaborative systems is not fully leveraged by offloading methods • Suitable methods for solving offloading problems depend on system complexity • Offloading methods are mainly evaluated using study-specific simulation environments • Studies lack cross-validation across methods, networks, and configurations The rapid increase in the number of connected devices has created a vast computational infrastructure. While intelligent applications (e.g., VR/AR, autonomous driving, and AI assistants) demand substantial processing power, they are typically executed only in short bursts. As a result, devices remain idle for significant periods, presenting an opportunity to utilize their unused computational resources. Computation offloading has emerged as a viable solution, where resource-constrained devices leverage edge or cloud infrastructure for intensive tasks. More recently, horizontal offloading has gained attention as a complementary approach, enabling devices within the same layer to collaborate and utilize idle resources more efficiently. In this survey, we examine current research on computation offloading in collaborative computing systems that span from horizontal offloading among end devices to more complex systems involving horizontal as well as vertical offloading to edge/cloud systems. Our examination focuses on architectures, collaborative characteristics, and offloading dynamics, such as task dependency, mobility, and multi-hop offloading. We begin by introducing edge computing paradigms and key concepts in computation offloading and collaborative computing. The offloading approaches are then classified into classical, heuristic, metaheuristic, and machine learning methods, further divided into centralized, decentralized, and distributed categories. These approaches are analyzed and compared based on collaborative system, architecture, offloading dynamics, and evaluation methodology. Finally, we have identified open challenges and emphasized the need to unify research across edge computing paradigms, alongside a focus on more complex collaborative systems that resemble real-world systems, and the development of standardized evaluation frameworks as key directions for future work.
Jussi Kalliola, Jason J. Jung, David Hästbacka
J. Parallel Distributed Comput.3
2026 A Systematic Mapping of federated learning operations and features: Architecture, communication and aggregation models
abstract
Federated Learning (FL) is a collaborative learning paradigm in which multiple clients train a shared global model without exchanging data. Clients communicate only model updates with a central aggregator. In parallel, Machine Learning Operations (MLOps) streamline the development, deployment, and monitoring of ML systems, while their extension, Federated Learning Operations (FLOps), aims to bring the operational discipline to decentralized and privacy-sensitive settings. This study presents a systematic mapping study (SMS) on FL and FLOps, and clarifies foundational concepts and uncovers new perspectives within this evolving field. We focus on FLOps and FL features: architecture, communication and aggregation models. First research question (RQ) focuses on prevalent FL computing architectures. Second RQ covers data transfer between FL components. Third RQ determines breadth of FLOps application in the scientific literature. Fourth RQ identifies distinct approaches to global model aggregation. Our analysis reveals that Edge-based local training with Cloud-based aggregation is the most adopted architecture, combining Edge privacy and responsiveness with Cloud computational capacity. Communication is enabled through lightweight protocols such as Message Queuing Telemetry Transport (MQTT), but protocol choice depends on constraints. Notably, FLOps remains a rarely addressed topic, indicating a substantial gap in end-to-end support for FL pipelines. Federated Averaging (FedAvg) is the most employed aggregation approach, valued for its simplicity and effectiveness with heterogeneous data. These findings expose critical research gaps in architectural diversity, protocol selection, lifecycle integration and adaptive aggregation, highlighting the need for more cohesive and scalable FL system design in future work.
Ari Kukkaro, Sergio Moreschini, Davide Taibi 0001, David Hästbacka
J. Syst. Softw.4
2025 AI Model Cards: State of the Art and Path to Automated Use
abstract
In software engineering, the integration of machine learning (ML) and artificial intelligence (AI) components into modern web services has become commonplace. To comply with evolving regulations, such as the EU AI Act, the development of AI models must adhere to the principles of transparency. This includes the training data used, the intended use, potential biases, and the risks associated with these models. To support these goals, documents named Model Cards were introduced to standardize ethical reporting and allow stakeholders to evaluate models based on various goals. In our ongoing research, we aim to automate risk analysis and regulatory compliance checks in software systems. We envision that model cards can serve as useful tools to achieve the goal. Given the evolving format of model cards over time, we conducted a state-of-the-art review of the current state and practice of model cards by analyzing 90 model cards from four model repositories to assess their relevance to our vision. The study's contribution is a thorough analysis of the model cards' structure and content, as well as their ethical reporting. Our study reveals the variance in information reporting, the loose structure, and the lack of ethical reporting in the model cards. Based on the findings, we propose a unified model card template that aims to enhance the structure, promote greater transparency, and establish a foundation for future machine-interpretable AI model cards.
Ali Mehraj, An Cao, Kari Systä, Tommi Mikkonen, Pyry Kotilainen, David Hästbacka, Niko Mäkitalo
WEBIST6
2024 Continuous Training vs. Transfer Learning on Edge and Fog Environments: A Steam Detection use Case
abstract
The implementation of smart manufacturing, which utilises advanced digital technologies to enhance the agility and productivity of the traditional manufacturing sector, has the potential to reduce resource consumption, optimise processes and enhance safety. One challenge in process automation (PA) is its strict real-time requirements. One solution to this challenge is the use of Edge and Fog computing platforms with finite computational power, which brings processing and data storing closer to the data sources. This proximity of computing devices reduces the latency and bandwidth requirements, relaxes the need for a reliable Internet connection, and provides more security in design over the Cloud solutions. This paper compares the performance of Edge and Fog computing for soft real-time machine learning-based visual process monitoring that supports the human operator. The objective is to get a better understanding how this ML task can be relocated within Edge and Fog layers. Moreover, the article provides con-siderations of emerging difficulties of practical implementation of Continuous Training pipeline and soft real-time steam detection.
Ari Kukkaro, Sergio Moreschini, David Hästbacka
SEAA3
2024 Best Practices for Resource Provisioning Declaration Within the Cognitive Cloud Continuum
abstract
The evolution of cloud computing, driven by ad-vances in mobile, edge technologies, and AI, has led to the development of the Cognitive Cloud Continuum (COCLCON). However, this paradigm introduces new challenges in managing and optimizing computing resources across a heterogeneous environment. This paper explores best practices for declaring resources within COCLCON, with a focus on efficient resource allocation and transparently declaring available resources by de-vices. In this study, we undertook a non-holistic literature review to identify current technologies used to specify requirements and to determine current gaps in best practices. The main outcome of our work is a proposed schema for Resource Provisioning Declaration, which will allow for increased knowledge related to the available devices and resources within the COCLCON.
Sergio Moreschini, Michele Albano, David Hästbacka
SEAA3
2024 The Pains and Gains of Microservices Revisited
Mikko Pirhonen, Kari Systä, David Hästbacka
PROFES3
2024 Edge to cloud tools: A Multivocal Literature Review
abstract
Edge-to-cloud computing is an emerging paradigm for distributing computational tasks between edge devices and cloud resources. Different approaches for orchestration, offloading, and many more purposes have been introduced in research. However, it is still not clear what has been implemented in the industry. This work aims to merge this gap by mapping the existing knowledge on edge-to-cloud tools by providing an overview of the current state of research in this area and identifying research gaps and challenges. For this purpose, we conducted a Multivocal Literature Review (MLR) by analyzing 40 tools from 1073 primary studies (220 PS from the white literature and 853 PS from the grey literature). We categorized the tools based on their characteristics and targeted environments. Overall, this systematic mapping study provides a comprehensive overview of edge-to-cloud tools and highlights several opportunities for researchers and practitioners for future research in this area. Editor’s note: Open Science material was validated by the Journal of Systems and Software Open Science Board.
Sergio Moreschini, Elham Younesian, David Hästbacka, Michele Albano, Jiri Hosek, Davide Taibi 0001
J. Syst. Softw.3
2023 Enhancing Collaborative Prototype Development: An Evaluation of the Descriptive Model for Prototyping Process
abstract
In this article, the ongoing research on collaborative prototype development between university and enterprises is presented. The study of project featured numerous pilot cases and prototypes, executed in collaboration with organizations to address real-world challenges. This article assesses the appropriateness of the Descriptive Model for Prototyping Process (DMPP) for research project applications. We delve into two primary facets: the synergy between universities and enterprises, and the potential for artifact reusability within the DMPP. The article presents various pilot cases from the KIEMI project, highlighting the DMPP’s role in each. Furthermore, the paper evaluates the model, sets forward the challenges faced, and, finally, discusses topics for future research.
Janne Harjamäki, Mika Saari, Mikko Nurminen, Petri Rantanen, Jari Soini, David Hästbacka
EJC6
2023 External Token-Based Authorization of Data-Driven Integrations and Service Compositions in MQTT 5
abstract
Modern connected cyber-physical systems and their integrations to traditional information systems are increasingly dependant on data and data sharing management in their integrations. Many such systems are constantly changing and evolving their composition, often including integrations to third party (data-driven) services. This paper presents a model where a service framework, used to manage microservice configurations, is also utilized to manage access to MQTT Version 5 message topics. A proof of concept is provided demonstrating how Eclipse Arrowhead as the service management layer is capable of taking care of authentication and authorization of publish and subscribe actions to MQTT topics as individually managed data services. The study shows that JSON Web Tokens (JWT) from this service framework can be used in the MQTT Version 5 headers without violating the MQTT specification as demonstrated with HiveMQ as the message broker in the proof of concept implementation.
David Hästbacka, Petri Kannisto, Mikael Filppula, Pál Varga
IECON1
2023 Open Data Platform Tools for Energy Service Ecosystem in Urban Superblocks
abstract
Superblocks, or large city-blocks with some degree of energy autonomy, have yet unanswered challenges related to data utilization. Within superblocks, a service-based data ecosystem could provide benefits in the form of higher-level control applications and flexibility for the energy community. In this paper, we use open data platforms and tools, namely FIWARE and Eclipse Arrowhead, to find a system design for energy data services. The research questions relate to recognizing what features and functions are required from data platforms to provide appropriate solution, and which of them are supported by the studied data platforms. The proposed solution is then tested with a prototype implementation in limited scale to gauge its viability. We conclude that FIWARE and Arrowhead complement each other and when used in unison fill most of the requirements established in this paper.
Mikael Filppula, Petri Kannisto, David Hästbacka
INDIN3
2023 Can We Trust the Default Vulnerabilities Severity?
abstract
As software systems become increasingly complex and interconnected, the risk of security debt has risen significantly, increasing cyber-attacks and data breaches. Vulnerability prioritization is a critical activity in software engineering as it helps identify and address security vulnerabilities in software systems promptly and effectively. With the increasing complexity of software systems and the growing number of potential threats, it is essential to have a systematic approach to vulnerability prioritization to ensure that the most critical vulnerabilities are addressed first. The present study aims to investigate the agreement between the default and the National Vulnerability Database (NVD) severity levels. We analyzed 1626 vulnerabilities encompassing 12 unique types of vulnerabilities associated with 125 Common Platform Enumeration identifiers belonging to 105 Apache projects. Our results show a scarce correlation between the default and NVD severity levels. Thus, the default severity of vulnerabilities is not trustworthy. Moreover, we discovered that, surprisingly, the same type of vulnerability has several NVD severity; therefore, no default prioritization can be accurate based only on the type of vulnerability. Future studies are needed to accurately estimate the priority of vulnerabilities by considering several aspects of vulnerabilities rather than only the type.
Matteo Esposito 0001, Sergio Moreschini, Valentina Lenarduzzi, David Hästbacka, Davide Falessi
SCAM4
2022 Data Autonomy in Message Brokers in Edge and Cloud for Mobile Machinery: Requirements and Technology Survey
abstract
The future data-driven manufacturing ecosystems build upon data spaces where each participant controls how its own data are utilized. This goal is equally important in machinery where the networks comprise machine fleets and the data use cases range from local edge to various cloud systems and digital business integrations. These systems of systems require efficient, scalable data streaming with decoupled (e.g., publish-subscribe) platforms that enable the flexible connection of data producers and consumers in heterogeneous networks. This paper describes a work in progress about data autonomy, sovereignty, and trust in message brokers in machinery, aiming to contribute to initiatives, such as Gaia-X and International Data Spaces (IDS). First, the paper identifies requirements for platforms and communication that span edge and cloud. Second, it presents a technology survey about data autonomy in open, Internet-cabable brokers. These include Advanced Message Queueing Protocol (AMQP), MQ Telemetry Transport (MQTT), Apache Kafka and Apache Pulsar. It appears that there is little research about data autonomy in brokers and MQTT has the strongest base. The work continues to develop data autonomy into a message broker.
Petri Kannisto, David Hästbacka
ETFA2
2022 Interoperability of OPC UA PubSub with Existing Message Broker Integration Architectures
abstract
Interoperable communication technologies are of key importance in production systems with increasing needs for data in their adoption of data-driven methodologies and new, emerging applications. OPC UA PubSub defines an alternative to the traditional client-server communication with a publish-subscribe model for data to cater to scalability and data-driven cloud application needs. In this paper, the OPC UA PubSub model is compared to some other message broker and communication technologies and integrated with an existing message based integration model for evaluating the interoperability. A case example is presented where data payloads and information security practices are integrated using an adapter approach.
David Hästbacka, Petri Kannisto, Antti Kätkytniemi
IECON1
2022 MLOps for evolvable AI intensive software systems
abstract
DevOps practices are the de facto sandard when developing software. The increased adoption of machine learning (ML) to solve problems urges us to adapt all the current approaches to developing a new standard that can take full benefit from the new solution. In this work we propose a graphical representation for DevOps for ML-based applications, namely MLOps, and also outline open research challenges. The pipeline aims to get the best of both worlds by maintaining the simple and iconic pipeline of DevOps, yet improving it by adding new circular steps for ML incorporation. This aims to create an ML-based development subsystem that can be self-maintained, and is capable of evolving side-by-side with the software development.
Sergio Moreschini, Francesco Lomio, David Hästbacka, Davide Taibi 0001
SANER3
2022 Dynamic Edge and Cloud Service Integration for Industrial IoT and Production Monitoring Applications of Industrial Cyber-Physical Systems
abstract
Industrial cyber-physical systems rely increasingly on data from Internet-of-Things (IoT) devices and other systems as continuously emerging use cases implement new intelligent features. Edge computing can be seen as an extension of the cloud in close physical proximity, in which some of the typical cloud computing loads are beneficial to run. This article studies data analytics application development for integration of industrial IoT data and composition of application services executed on edge and cloud. A solution is designed to support heterogeneous hardware and run-time platforms, and focuses on the service layer that enables flexible orchestration of data flows and dynamic service compositions. The unified model and system architecture implemented, using the open Arrowhead framework model, is verified through two representative industrial use cases.
David Hästbacka, Jari Halme, Laurentiu Barna, Henrikki Hoikka, Henri Pettinen, Martin Larrañaga, Mikael Björkbom, Heikki Mesiä, Antti Jaatinen, Marko Elo
IEEE Trans. Ind. Informatics1
2019 Monitoring of Production Processes and the Condition of the Production Equipment through the Internet
abstract
The decreasing prices of monitoring equipment have vastly increased the opportunities to utilize local data, and data processing for wider global web-based monitoring purposes. The possible amount of data flowing though different levels can be huge. Now the question is how to handle this opportunity in both dynamic and secure way. The paper presents a new concept to manage data for monitoring through the Internet. The concept is based on the use of Arrowhead Framework (AF) and MIMOSA data model, and selected edge, and gateway devices together with cloud computing opportunities. The concept enables the flexible and secure orchestration of run-time data sources and the utilization of computational services for various process and condition monitoring needs.
Jari Halme, Erkki Jantunen, David Hästbacka, Csaba Hegedüs, Pál Varga, Mikael Björkbom, Heikki Mesiä, Rupesh More, Antti Jaatinen, Laurentiu Barna, Pasi Tuominen, Henri Pettinen, Marko Elo, Martin Larrañaga
CoDIT3
2018 Data-driven and Event-driven Integration Architecture for Plant-wide Industrial Process Monitoring and Control
abstract
Efficiency of industrial processes and a high quality of the products can be achieved with advanced monitoring and control solutions. In addition, industrial processes often consume large amounts of energy and through their efficiency and use of resources also have a significant environmental impact. For industrial process optimisation it is necessary to integrate distributed data and functionality into plant-wide coordinating level solutions. From an implementation point of view this is challenging due to different communication protocols and messaging structures. In this paper an integration architecture is proposed that decouples control systems using a message bus mediator approach. The mediator acts as a unified point of access that through adapters facilitates integration of existing control systems to advanced plant-wide data-driven and event-driven control. It is demonstrated with a laboratory case presenting two examples applicable to real-life industrial problems.
David Hästbacka, Petri Kannisto, Matti Vilkko
IECON1
2016 Context modeling with situation rules for industrial maintenance
abstract
Industrial maintenance requires not only experienced service personnel to carry out the tasks but also up-to-date information about the target equipment and its environment. Accessing information required to execute the tasks is a common challenge for maintenance personnel. This paper presents a knowledge modeling approach and a technical architecture of a gateway system developed to support maintenance personnel with information combined from legacy data sources as well as from context ontology augmented with situational knowledge. The novelty of the approach is its unified object oriented style of knowledge representation encapsulating predefined queries and rules into ontology classes. The approach utilizes standard Semantic Web technologies, especially SPARQL query language and SPARQL Inferencing Notation SPIN. Feasibility of the approach is demonstrated with a simple maintenance use case example executed in an experimental knowledge gateway system.
Pekka Aarnio, Valeriy Vyatkin, David Hästbacka
ETFA3
2016 Cloud-Based Management of Machine Learning Generated Knowledge for Fleet Data Refinement
Petri Kannisto, David Hästbacka
IC3K2
2016 Service-based condition monitoring for cloud-enabled maintenance operations
abstract
Condition based maintenance is regarded as a maintenance strategy that through measured component wear balances availability and accurate operation with necessary maintenance operations. In industrial settings there can be hundreds or thousands of objects to monitor, and systems that are distributed into different networks are seldom compatible. This paper proposes a condition monitoring system based on standardized services operating as parts of a service framework. The solution builds on dynamic service composition as well as standard information models for measurement related data, and supports functionality from sensors to cloud applications. The approach has been applied to industrial processing in mining and it can be claimed to improve interoperability as well as reduce engineering effort when composing functionality supporting maintenance operations.
David Hästbacka, Erkki Jantunen, Mika Karaila, Laurentiu Barna
IECON1
2014 Device status information service architecture for condition monitoring using OPC UA
abstract
Condition monitoring and maintenance of devices and equipment is an important aspect of operating a production facility affecting the availability of production systems. Modern production environments can consist of thousands of devices that each need to be monitored so that maintenance can be performed when necessary to sustain a cost-effective state of production. Today operation and maintenance (O&M) is typically outsourced, and equipment and device manufacturers have also entered the service business. This brings challenges in managing a multitude of different devices using different protocols as well as in the varying needs for utilizing this information in enterprise functions and services. Based on OPC Unified Architecture (UA) a scalable architecture is developed for providing device status information of heterogeneous field devices and sensors to enterprise level applications and services. A proof of concept implementation of this architecture is presented and its envisioned adoption in a mine environment is discussed.
David Hästbacka, Laurentiu Barna, Mika Karaila, Yiqing Liang, Pasi Tuominen, Seppo Kuikka
ETFA1
2013 Semantics enhanced engineering and model reasoning for control application development
David Hästbacka, Seppo Kuikka
Multim. Tools Appl.1
2012 Facilitating services and engineering process management in distributed engineering of control applications
abstract
Engineering of information and control systems for process and manufacturing facilities is a complex multidisciplinary effort. The engineering processes are often interwoven between different teams and enterprises making both the exchange of technical engineering data and management of engineering processes challenging. Business process management is applied to the engineering processes in order to improve engineering effectiveness and ease management of distributed activities. Modeling engineering processes as business processes and implementing them executable enables many tasks to be automated and supported with information systems. In combination with service bus based system integration the information exchange and the use of various services can be enhanced. A prototype service infrastructure with engineering services has been implemented to prove the concept. Exemplary engineering processes have been developed to demonstrate the feasibility of executing routine tasks automatically and improving management of both human and information system tasks.
David Hästbacka, Seppo Kuikka
INDIN1
2011 Model-driven development of industrial process control applications
David Hästbacka, Timo Vepsäläinen, Seppo Kuikka
J. Syst. Softw.1
2008 Tool Support for the UML Automation Profile - For Domain-Specific Software Development in Manufacturing
abstract
The development of modern distributed automation applications is challenging and present development practices contain manual transferring of informal information from one phase to another. Our research aims to overcome some of these challenges by integrating concepts from modern object-oriented design, model-driven development and high-level modeling potential of the UML automation profile into a seamless development path from PI-diagrams to control software. This paper presents a prototype of a control engineering tool that supports the UML automation profile and is intended to cover part of the development chain. The tool was implemented on the Eclipse platform and it utilizes various open source tools and frameworks to enable also usage of UML and SysML in modeling work. The implemented tool can be extended by transformation tools capable of processing requirements of the control system and PIM-model of the designed control software.
Timo Vepsäläinen, David Hästbacka, Seppo Kuikka
ICSEA2