VLDB 2026 Research / reviewers in the wild / expert
Firas Bayram
dblp:283/7067
· DBLP profile ↗
8ranked-venue papers
7as first author
8since 2021 · last 2024
0000-0003-0683-2783ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 4 · 3 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Adaptive data quality scoring operations framework using drift-aware mechanism for industrial applicationsabstractWithin data-driven artificial intelligence (AI) systems for industrial applications, ensuring the reliability of theincoming data streams is an integral part of trustworthy decision-making. An approach to assess data validityis data quality scoring, which assigns a score to each data point or stream based on various quality dimensions.However, certain dimensions exhibit dynamic qualities, which require adaptation on the basis of the system’scurrent conditions. Existing methods often overlook this aspect, making them inefficient in dynamic productionenvironments. In this paper, we introduce the Adaptive Data Quality Scoring Operations Framework, a novelframework developed to address the challenges posed by dynamic quality dimensions in industrial data streams.The framework introduces an innovative approach by integrating a dynamic change detector mechanism thatactively monitors and adapts to changes in data quality, ensuring the relevance of quality scores. We evaluatethe proposed framework performance in a real-world industrial use case. The experimental results reveal highpredictive performance and efficient processing time, highlighting its effectiveness in practical quality-drivenAI applications. Firas Bayram, Bestoun S. Ahmed, Erik Hallin |
J. Syst. Softw. | 1 |
| 2023 | DQSOps: Data Quality Scoring Operations Framework for Data-Driven ApplicationsabstractData quality assessment has become a prominent component in the successful execution of complex data-driven artificial intelligence (AI) software systems. In practice, real-world applications generate huge volumes of data at speeds. These data streams require analysis and preprocessing before being permanently stored or used in a learning task. Therefore, significant attention has been paid to the systematic management and construction of high-quality datasets. Nevertheless, managing voluminous and high-velocity data streams is usually performed manually (i.e. offline), making it an impractical strategy in production environments. To address this challenge, DataOps has emerged to achieve life-cycle automation of data processes using DevOps principles. However, determining the data quality based on a fitness scale constitutes a complex task within the framework of DataOps. This paper presents a novel Data Quality Scoring Operations (DQSOps) framework that yields a quality score for production data in DataOps workflows. The framework incorporates two scoring approaches, an ML prediction-based approach that predicts the data quality score and a standard-based approach that periodically produces the ground-truth scores based on assessing several data quality dimensions. We deploy the DQSOps framework in a real-world industrial use case. The results show that DQSOps achieves significant computational speedup rates compared to the conventional approach of data quality scoring while maintaining high prediction performance. Firas Bayram, Bestoun S. Ahmed, Erik Hallin, Anton Engman |
EASE | 1 |
| 2023 | DA-LSTM: A dynamic drift-adaptive learning framework for interval load forecasting with LSTM networksabstractLoad forecasting is a crucial topic in energy management systems (EMS) due to its vital role in optimizing energy scheduling and enabling more flexible and intelligent power grid systems. As a result, these systems allow power utility companies to respond promptly to demands in the electricity market. Deep learning (DL) models have been commonly employed in load forecasting problems supported by adaptation mechanisms to cope with the changing pattern of consumption by customers, known as concept drift. A drift magnitude threshold should be defined to design change detection methods to identify drifts. While the drift magnitude in load forecasting problems can vary significantly over time, existing literature often assumes a fixed drift magnitude threshold, which should be dynamically adjusted rather than fixed during system evolution. To address this gap, in this paper, we propose a dynamic drift-adaptive Long Short-Term Memory (DA-LSTM) framework that can improve the performance of load forecasting models without requiring a drift threshold setting. We integrate several strategies into the framework based on active and passive adaptation approaches. To evaluate DA-LSTM in real-life settings, we thoroughly analyze the proposed framework and deploy it in a real-world problem through a cloud-based environment. Efficiency is evaluated in terms of the prediction performance of each approach and computational cost. The experiments show performance improvements on multiple evaluation metrics achieved by our framework compared to baseline methods from the literature. Finally, we present a trade-off analysis between prediction performance and computational costs. Firas Bayram, Phil Aupke, Bestoun S. Ahmed, Andreas Kassler, Andreas Theocharis 0001, Jonas Forsman |
Eng. Appl. Artif. Intell. | 1 |
| 2023 | A domain-region based evaluation of ML performance robustness to covariate shiftabstractAbstract Most machine learning methods assume that the input data distribution is the same in the training and testing phases. However, in practice, this stationarity is usually not met and the distribution of inputs differs, leading to unexpected performance of the learned model in deployment. The issue in which the training and test data inputs follow different probability distributions while the input–output relationship remains unchanged is referred to as covariate shift. In this paper, the performance of conventional machine learning models was experimentally evaluated in the presence of covariate shift. Furthermore, a region-based evaluation was performed by decomposing the domain of probability density function of the input data to assess the classifier’s performance per domain region. Distributional changes were simulated in a two-dimensional classification problem. Subsequently, a higher four-dimensional experiments were conducted. Based on the experimental analysis, the Random Forests algorithm is the most robust classifier in the two-dimensional case, showing the lowest degradation rate for accuracy and F1-score metrics, with a range between 0.1% and 2.08%. Moreover, the results reveal that in higher-dimensional experiments, the performance of the models is predominantly influenced by the complexity of the classification function, leading to degradation rates exceeding 25% in most cases. It is also concluded that the models exhibit high bias toward the region with high density in the input space domain of the training samples. Firas Bayram, Bestoun S. Ahmed |
Neural Comput. Appl. | 1 |
| 2022 | A Drift Handling Approach for Self-Adaptive ML Software in Scalable Industrial ProcessesabstractMost industrial processes in real-world manufacturing applications are characterized by the scalability property, which requires an automated strategy to self-adapt machine learning (ML) software systems to the new conditions. In this paper, we investigate an Electroslag Remelting (ESR) use case process from the Uddeholms AB steel company. The use case involves predicting the minimum pressure value for a vacuum pumping event. Taking into account the long time required to collect new records and efficiently integrate the new machines with the built ML software system. Additionally, to accommodate the changes and satisfy the non-functional requirement of the software system, namely adaptability, we propose an automated and adaptive approach based on a drift handling technique called importance weighting. The aim is to address the problem of adding a new furnace to production and enable the adaptability attribute of the ML software. The overall results demonstrate the improvements in ML software performance achieved by implementing the proposed approach over the classical non-adaptive approach. Firas Bayram, Bestoun S. Ahmed, Erik Hallin, Anton Engman |
ASE | 1 |
| 2022 | A systematic mapping study of source code representation for deep learning in software engineeringabstractAbstract The usage of deep learning (DL) approaches for software engineering has attracted much attention, particularly in source code modelling and analysis. However, in order to use DL, source code needs to be formatted to fit the expected input form of DL models. This problem is known as source code representation. Source code can be represented via different approaches, most importantly, the tree‐based, token‐based, and graph‐based approaches. We use a systematic mapping study to investigate i detail the representation approaches adopted in 103 studies that use DL in the context of software engineering. Thus, studies are collected from 2014 to 2021 from 14 different journals and 27 conferences. We show that each way of representing source code can provide a different, yet orthogonal view of the same source code. Thus, different software engineering tasks might require different (combinations of) code representation approaches, depending on the nature and complexity of the task. Particularly, we show that it is crucial to define whether the DL approach requires lexical, syntactical, or semantic code information. Our analysis shows that a wide range of different representations and combinations of representations (hybrid representations) are used to solve a wide range of common software engineering problems. However, we also observe that current research does not generally attempt to transfer existing representations or models to other studies even though there are other contexts in which these representations and models may also be useful. We believe that there is potential for more reuse and the application of transfer learning when applying DL to software engineering tasks. Hazem Peter Samoaa, Firas Bayram, Pasquale Salza, Philipp Leitner 0001 |
IET Softw. | 2 |
| 2022 | From concept drift to model degradation: An overview on performance-aware drift detectorsabstractThe dynamicity of real-world systems poses a significant challenge to deployed predictive machine learning (ML) models. Changes in the system on which the ML model has been trained may lead to performance degradation during the system’s life cycle. Recent advances that study non-stationary environments have mainly focused on identifying and addressing such changes caused by a phenomenon called concept drift. Different terms have been used in the literature to refer to the same type of concept drift and the same term for various types. This lack of unified terminology is set out to create confusion on distinguishing between different concept drift variants. In this paper, we start by grouping concept drift types by their mathematical definitions and survey the different terms used in the literature to build a consolidated taxonomy of the field. We also review and classify performance-based concept drift detection methods proposed in the last decade. These methods utilize the predictive model’s performance degradation to signal substantial changes in the systems. The classification is outlined in a hierarchical diagram to provide an orderly navigation between the methods. We present a comprehensive analysis of the main attributes and strategies for tracking and evaluating the model’s performance in the predictive system. The paper concludes by discussing open research challenges and possible research directions. Firas Bayram, Bestoun S. Ahmed, Andreas Kassler |
Knowl. Based Syst. | 1 |
| 2021 | Predicting Tennis Match Outcomes with Network Analysis and Machine Learning
Firas Bayram, Davide Garbarino, Annalisa Barla |
SOFSEM | 1 |