Maroua Bahri

dblp:201/3890 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
10since 2021 · last 2025
0000-0002-7420-7464ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 9 first-author · 7 since 2021Databases, data management, data science and information retrieval · 9 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorTheory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 OSMAC: A Dynamic SMAC for Data Streams
abstract
Automated machine learning (autoML) methods often require multiple passes over data and are computationally intensive, rendering them unsuitable for streaming scenarios where data is continuously generated and distributions evolve over time. The few existing autoML solutions for stream learning mainly rely on random search or genetic algorithms, which struggle to maintain high performance in dynamic environments. By contrast, leading methods in batch learning such as the Sequential Model-based Algorithm Configuration (SMAC) leverage modelbased approaches, suggesting opportunities for improvement in stream settings. To address these challenges and meet the requirements of stream scenarios, we introduce OnlineSMAC, a model-based optimizer for data streams. OnlineSMAC combines Bayesian optimization with an extension of the SMAC optimizer to dynamically select optimal processing pipelines and hyperparameters. Our results show that this approach is highly competitive, achieving performance on par with state-of-the-art stream autoML methods. This highlights the promising potential of using Bayesian optimization for data streams.
Émile Royer, Maroua Bahri, Nikolaos Georgantas
ICTAI2
2025 Bayesian Stream Tuner: Dynamic Hyperparameter Optimization for Real-Time Data Streams
abstract
Hyperparameter optimization is crucial for maximizing machine learning model performance, yet most existing algorithms are designed for batch or offline scenarios and assume static data distributions.Such assumptions fall short in data stream settings, where models must adapt to evolving inputs in real time.To address these limitations, we propose the Bayesian Stream Tuner (BST), a novel framework for online hyperparameter optimization in nonstationary data streams.BST maintains a dynamic set of candidate hyperparameter configurations and periodically refines them using an incremental Bayesian model, which estimates configuration performance based on recent data statistics and hyperparameter values.This systematic exploration and refinement strategy allows BST to detect and respond to concept drift by resetting its adaptation mechanisms whenever necessary, ensuring strong performance under changing distributions.Our theoretical analysis establishes sublinear regret bounds for BST in dynamic environments, and extensive experiments on classification and regression tasks demonstrate that BST consistently outperforms state-of-the-art online hyperparameter optimization methods in both predictive accuracy and adaptability, making it a powerful solution for real-time hyperparameter tuning in evolving data streams.
Nilesh Verma, Albert Bifet, Bernhard Pfahringer, Maroua Bahri
KDD (2)4
2025 Auto-Reg: A Dynamic AutoML Framework for Streaming Regression
Nilesh Verma, Albert Bifet, Bernhard Pfahringer, Maroua Bahri
PAKDD (4)4
2025 AutoSAD: An Adaptive Framework for Streaming Anomaly Detection
Nilesh Verma, Albert Bifet, Bernhard Pfahringer, Maroua Bahri
PRICAI (5)4
2023 AutoClass: AutoML for Data Stream Classification
abstract
Automated Machine Learning (autoML) is a novel topic that aims to tackle the parameter configuration issue using automatic monitoring models and comprises different machine learning tasks, such as feature selection, model selection, and hyper-parameter tuning. It makes easier use of algorithms for non-ML experts as well as ML experts by automating tasks that rely on expert domain knowledge. Nevertheless, autoML is in its infancy stage and not well explored yet in the offline and stream settings. In this paper, we propose automated Classification (auto-Class) method for automated algorithm selection and configuration for data stream classification. AutoClass consists of training an ensemble of different tuned configurations and selecting the best-performing configuration to do the prediction. We present experiments performed on a diverse set of real and artificial datasets and show how our proposed approach can outperform the performance of competitive state-of-the-art ensemble and single-based methods.
Maroua Bahri, Nikolaos Georgantas
IEEE Big Data1
2022 STREamRHF: Tree-Based Unsupervised Anomaly Detection for Data Streams
abstract
We present STREAMRHF, an unsupervised anomaly detection algorithm for data streams. Our algorithm builds on some of the ideas of Random Histogram Forest (RHF) [1], a state-of-the-art algorithm for batch unsupervised anomaly detection. STREAMRHF constructs a forest of decision trees, where feature splits are determined according to the kurtosis score of every feature. It irrevocably assigns an anomaly score to data points, as soon as they arrive, by means of an incremental computation of its random trees and the kurtosis scores of the features. This allows efficient online scoring and concept drift detection altogether. Our approach is tree-based which boasts several appealing properties, such as explainability of the results [2]. We conduct an extensive experimental evaluation on multiple datasets from different real-world applications. Our evaluation shows that our streaming algorithm achieves comparable average precision to RHF while outperforming state-of-the-art streaming approaches for unsupervised anomaly detection with furthermore limited computational complexity.
Stefan Nesic, Andrian Putina, Maroua Bahri, Alexis Huet, José Manuel Navarro, Dario Rossi 0001, Mauro Sozio
AICCSA3
2022 Effective Weighted k-Nearest Neighbors for Dynamic Data Streams
abstract
Many real-world applications involve classification from evolving data streams. However, learning in such environment requires algorithms able to learn and predict from potentially unbounded data that are constantly changing. For this to happen, stream algorithms should restrict the storage to a part of – and/or synopsis information from – the stream using efficient and accurate manners and strategies, such as window models and summarization techniques (e.g., sampling, sketching, dimensionality reduction). In this work, we focus on the k-Nearest Neighbors (kNN) where most of the existing approaches for data streams consider that instances have the same weight from the start to the finish of the processing task.In a streaming data scenario, it is often the case that the most recent elements from the data stream are the more relevant ones. Taking into account that the most recent instances are more relevant, we propose a novel kNN approach that stores instances in a sliding window and weighs them according to their arrival time (i.e position on the window) using an adjusted weight function. The empirical results on comprehensive real and synthetic datasets indicate the effectiveness and efficiency of our proposed approach in comparison with state-of-the-art algorithms.
Maroua Bahri
IEEE Big Data1
2022 AutoAD: an Automated Framework for Unsupervised Anomaly Detectio
abstract
Over the last decade, we witnessed the proliferation of several machine learning algorithms capable of solving different tasks for the most diverse applications. Often, for an algorithm to be effective, significant human effort is required, in particular for hyper-parameter tuning and data cleaning. Recently, there have been increasing efforts to alleviate such a burden and make machine learning algorithms easier to use for researchers with varying levels of expertise. Nevertheless, the question of whether an efficient and fully generalizable automated Machine Learning (autoML) framework is possible remains unanswered. In this paper, we present autoAD, the first autoML framework for unsupervised anomaly detection. By leveraging a pool of different anomaly detection algorithms, each one coming with its own hyper-parameter search space, our framework automatically selects the best performing approach, while determining an optimal configuration for its hyper-parameters on a given dataset. Our extensive experimental evaluation, conducted on a rich collection of datasets, shows the substantial gains that can be achieved with autoAD compared to state-of-the-art methods for unsupervised anomaly detection.
Andrian Putina, Maroua Bahri, Flavia Salutari, Mauro Sozio
DSAA2
2022 Evolution-Based Online Automated Machine Learning
Cedric Kulbach, Jacob Montiel, Maroua Bahri, Marco Heyden, Albert Bifet
PAKDD (1)3
2021 Incremental k-Nearest Neighbors Using Reservoir Sampling for Data Streams
Maroua Bahri, Albert Bifet
DS1
2020 AutoML for Stream k-Nearest Neighbors Classification
abstract
The last few decades have witnessed a significant evolution of technology in different domains, changing the way the world operates, which leads to an overwhelming amount of data generated in an open-ended way as streams. Over the past years, we observed the development of several machine learning algorithms to process big data streams. However, the accuracy of these algorithms is very sensitive to their hyper-parameters, which requires expertise and extensive trials to tune. Another relevant aspect is the high-dimensionality of data, which can causes degradation to computational performance. To cope with these issues, this paper proposes a stream k-nearest neighbors (kNN) algorithm that applies an internal dimension reduction to the stream in order to reduce the resource usage and uses an automatic monitoring system that tunes dynamically the configuration of the kNN algorithm and the output dimension size with big data streams. Experiments over a wide range of datasets show that the predictive and computational performances of the kNN algorithm are improved.
Maroua Bahri, Bruno M. Veloso, Albert Bifet, João Gama 0001
IEEE BigData1
2020 Compressed k-Nearest Neighbors Ensembles for Evolving Data Streams
abstract
International audience
Maroua Bahri, Albert Bifet, Silviu Maniu, Rodrigo Fernandes de Mello, Nikolaos Tziortziotis
ECAI1
2020 Efficient Batch-Incremental Classification Using UMAP for Evolving Data Streams
abstract
Learning from potentially infinite and high-dimensional data streams poses significant challenges in the classification task. For instance, k -Nearest Neighbors ( k NN) is one of the most often used algorithms in the data stream mining area that proved to be very resource-intensive when dealing with high-dimensional spaces. Uniform Manifold Approximation and Projection (UMAP) is a novel manifold technique and one of the most promising dimension reduction and visualization techniques in the non-streaming setting because of its high performance in comparison with competitors. However, there is no version of UMAP that copes with the challenging context of streams. To overcome these restrictions, we propose a batch-incremental approach that pre-processes data streams using UMAP, by producing successive embeddings on a stream of disjoint batches in order to support an incremental k NN classification. Experiments conducted on publicly available synthetic and real-world datasets demonstrate the substantial gains that can be achieved with our proposal compared to state-of-the-art techniques.
Maroua Bahri, Bernhard Pfahringer, Albert Bifet, Silviu Maniu
IDA1
2020 Survey on Feature Transformation Techniques for Data Streams
abstract
Mining high-dimensional data streams poses a fundamental challenge to machine learning as the presence of high numbers of attributes can remarkably degrade any mining task's performance. In the past several years, dimension reduction (DR) approaches have been successfully applied for different purposes (e.g., visualization). Due to their high-computational costs and numerous passes over large data, these approaches pose a hindrance when processing infinite data streams that are potentially high-dimensional. The latter increases the resource-usage of algorithms that could suffer from the curse of dimensionality. To cope with these issues, some techniques for incremental DR have been proposed. In this paper, we provide a survey on reduction approaches designed to handle data streams and highlight the key benefits of using these approaches for stream mining algorithms.
Maroua Bahri, Albert Bifet, Silviu Maniu, Heitor Murilo Gomes
IJCAI1
2020 CS-ARF: Compressed Adaptive Random Forests for Evolving Data Stream Classification
abstract
Ensemble-based methods are one of the most often used methods in the classification task that have been adapted to the stream setting because of their high learning performance achievement. For instance, Adaptive Random Forests (ARF) is a recent ensemble method for evolving data streams that proved to be of a good predictive performance but, as all ensemble methods, it suffers from a severe drawback related to the high computational demand which prevents it from being efficient and further exacerbates with high-dimensional data. In this context, the application of a dimensionality reduction technique is crucial while processing the Internet of Things (IoT) data stream with ultrahigh dimensionality. In this paper, we aim to alleviate this deficiency and improve ARF performance, so we introduce the CS-ARF approach that uses Compressed Sensing (CS) as an internal pre-processing task, to reduce the dimensionality of data before starting the learning process, that will potentially lead to a meaningful improvement in memory usage. Experiments on various datasets show the high classification performance of our CS-ARF approach compared against current state-of-the-art methods while reducing resource usage.
Maroua Bahri, Heitor Murilo Gomes, Albert Bifet, Silviu Maniu
IJCNN1
2018 A Sketch-Based Naive Bayes Algorithms for Evolving Data Streams
abstract
A well-known learning task in big data stream mining is classification. Extensively studied in the offline setting, in the streaming setting - where data are evolving and even infinite - it is still a challenge. In the offline setting, training needs to store all the data in memory for the learning task; yet, in the streaming setting, this is impossible to do due to the massive amount of data that is generated in real-time. To cope with these resource issues, this paper proposes and analyzes several evolving naive Bayes classification algorithms, based on the well-known count-min sketch, in order to minimize the space needed to store the training data. The proposed algorithms also adapt concept drift approaches, such as ADWIN, to deal with the fact that streaming data may be evolving and change over time. However, handling sparse, very high-dimensional data in such framework is highly challenging. Therefore, we include the hashing trick, a technique for dimensionality reduction, to compress that down to a lower dimensional space, which leads to a large memory saving.We give a theoretical analysis which demonstrates that our proposed algorithms provide a similar accuracy quality to the classical big data stream mining algorithms using a reasonable amount of resources. We validate these theoretical results by an extensive evaluation on both synthetic and real-world datasets.
Maroua Bahri, Silviu Maniu, Albert Bifet
IEEE BigData1
2016 Clustering data stream under a belief function framework
abstract
Clustering is a crucial task for massive data that continuously arrive and evolve over time, generated as stream. However, data may be pervaded by uncertainty and imprecision, and techniques that achieve the unsupervised learning with imperfect data sets are unable to deal with such evolving environment. On the other hand, standard methods for clustering data streams are not adapted to an uncertain framework. Hence, in this paper, we propose a method for clustering data stream in an imperfect context, particularly using belief function theory in order to handle the belonging of objects to singletons and disjunctions of clusters.
Maroua Bahri, Zied Elouedi
AICCSA1