VLDB 2026 Research / reviewers in the wild / expert
Erik Hallin
dblp:326/2231
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2024
0009-0008-8052-6577ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 4 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Adaptive data quality scoring operations framework using drift-aware mechanism for industrial applicationsabstractWithin data-driven artificial intelligence (AI) systems for industrial applications, ensuring the reliability of theincoming data streams is an integral part of trustworthy decision-making. An approach to assess data validityis data quality scoring, which assigns a score to each data point or stream based on various quality dimensions.However, certain dimensions exhibit dynamic qualities, which require adaptation on the basis of the system’scurrent conditions. Existing methods often overlook this aspect, making them inefficient in dynamic productionenvironments. In this paper, we introduce the Adaptive Data Quality Scoring Operations Framework, a novelframework developed to address the challenges posed by dynamic quality dimensions in industrial data streams.The framework introduces an innovative approach by integrating a dynamic change detector mechanism thatactively monitors and adapts to changes in data quality, ensuring the relevance of quality scores. We evaluatethe proposed framework performance in a real-world industrial use case. The experimental results reveal highpredictive performance and efficient processing time, highlighting its effectiveness in practical quality-drivenAI applications. Firas Bayram, Bestoun S. Ahmed, Erik Hallin |
J. Syst. Softw. | 3 |
| 2023 | DQSOps: Data Quality Scoring Operations Framework for Data-Driven ApplicationsabstractData quality assessment has become a prominent component in the successful execution of complex data-driven artificial intelligence (AI) software systems. In practice, real-world applications generate huge volumes of data at speeds. These data streams require analysis and preprocessing before being permanently stored or used in a learning task. Therefore, significant attention has been paid to the systematic management and construction of high-quality datasets. Nevertheless, managing voluminous and high-velocity data streams is usually performed manually (i.e. offline), making it an impractical strategy in production environments. To address this challenge, DataOps has emerged to achieve life-cycle automation of data processes using DevOps principles. However, determining the data quality based on a fitness scale constitutes a complex task within the framework of DataOps. This paper presents a novel Data Quality Scoring Operations (DQSOps) framework that yields a quality score for production data in DataOps workflows. The framework incorporates two scoring approaches, an ML prediction-based approach that predicts the data quality score and a standard-based approach that periodically produces the ground-truth scores based on assessing several data quality dimensions. We deploy the DQSOps framework in a real-world industrial use case. The results show that DQSOps achieves significant computational speedup rates compared to the conventional approach of data quality scoring while maintaining high prediction performance. Firas Bayram, Bestoun S. Ahmed, Erik Hallin, Anton Engman |
EASE | 3 |
| 2022 | A Drift Handling Approach for Self-Adaptive ML Software in Scalable Industrial ProcessesabstractMost industrial processes in real-world manufacturing applications are characterized by the scalability property, which requires an automated strategy to self-adapt machine learning (ML) software systems to the new conditions. In this paper, we investigate an Electroslag Remelting (ESR) use case process from the Uddeholms AB steel company. The use case involves predicting the minimum pressure value for a vacuum pumping event. Taking into account the long time required to collect new records and efficiently integrate the new machines with the built ML software system. Additionally, to accommodate the changes and satisfy the non-functional requirement of the software system, namely adaptability, we propose an automated and adaptive approach based on a drift handling technique called importance weighting. The aim is to address the problem of adding a new furnace to production and enable the adaptability attribute of the ML software. The overall results demonstrate the improvements in ML software performance achieved by implementing the proposed approach over the classical non-adaptive approach. Firas Bayram, Bestoun S. Ahmed, Erik Hallin, Anton Engman |
ASE | 3 |
| 2022 | Testing of machine learning models with limited samples: an industrial vacuum pumping applicationabstractThere is often a scarcity of training data for machine learning (ML) classification and regression models in industrial production, especially for time-consuming or sparsely run manufacturing processes. Traditionally, a majority of the limited ground-truth data is used for training, while a handful of samples are left for testing. In that case, the number of test samples is inadequate to properly evaluate the robustness of the ML models under test (i.e., the system under test) for classification and regression. Furthermore, the output of these ML models may be inaccurate or even fail if the input data differ from the expected. This is the case for ML models used in the Electroslag Remelting (ESR) process in the refined steel industry to predict the pressure in a vacuum chamber. A vacuum pumping event that occurs once a workday generates a few hundred samples in a year of pumping for training and testing. In the absence of adequate training and test samples, this paper first presents a method to generate a fresh set of augmented samples based on vacuum pumping principles. Based on the generated augmented samples, three test scenarios and one test oracle are presented to assess the robustness of an ML model used for production on an industrial scale. Experiments are conducted with real industrial production data obtained from Uddeholms AB steel company. The evaluations indicate that Ensemble and Neural Network are the most robust when trained on augmented data using the proposed testing strategy. The evaluation also demonstrates the proposed method's effectiveness in checking and improving ML algorithms' robustness in such situations. The work improves software testing's state-of-the-art robustness testing in similar settings. Finally, the paper presents an MLOps implementation of the proposed approach for real-time ML model prediction and action on the edge node and automated continuous delivery of ML software from the cloud. Bestoun S. Ahmed, Erik Hallin, Anton Engman |
ESEC/SIGSOFT FSE | 3 |