Sedir Mohammed

dblp:357/3188 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
0000-0002-2163-0437ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 4 · 3 first-author · 4 since 2021
YearPublicationVenuePosition
2025 Data Quality in the Age of AI
Felix Naumann, Lisa Ehrlinger, Hazar Harmouch, Sedir Mohammed, Divesh Srivastava
ADBIS4
2025 Step-by-Step Data Cleaning Recommendations to Improve ML Prediction Accuracy
abstract
Data quality is crucial in machine learning (ML) applications, as errors in the data can significantly impact the prediction accuracy of the underlying ML model. Therefore, data cleaning is an integral component of any ML pipeline. However, in practical scenarios, data cleaning incurs significant costs, as it often involves domain experts for configuring and executing the cleaning process. Thus, efficient resource allocation during data cleaning can enhance ML prediction accuracy while controlling expenses. This paper presents COMET, a system designed to optimize data cleaning efforts for ML tasks. COMET gives step-by-step recommendations on which feature to clean next, maximizing the efficiency of data cleaning under resource constraints. We evaluated COMET across various datasets, ML algorithms, and data error types, demonstrating its robustness and adaptability. Our results show that COMET consistently outperforms feature importance-based, random, and another well-known cleaning method, achieving up to 52 and on average 5 percentage points higher ML prediction accuracy than the proposed baselines.
Sedir Mohammed, Felix Naumann, Hazar Harmouch
EDBT1
2025 The effects of data quality on machine learning performance on tabular data
abstract
Modern artificial intelligence (AI) applications require large quantities of training and test data. This need creates critical challenges not only concerning the availability of such data, but also regarding its quality. For example, incomplete, erroneous, or inappropriate training data can lead to unreliable models that produce ultimately poor decisions. Trustworthy AI applications require high-quality training and test data along many quality dimensions, such as accuracy, completeness, and consistency. We explore empirically the relationship between six data quality dimensions and the performance of 19 popular machine learning algorithms covering the tasks of classification, regression, and clustering, with the goal of explaining their performance in terms of data quality. Our experiments distinguish three scenarios based on the AI pipeline steps that were fed with polluted data: polluted training data, test data, or both. We conclude the paper with an extensive discussion of our observations.
Sedir Mohammed, Lukas Budach, Moritz Feuerpfeil, Nina Ihde, Andrea Nathansen, Nele Sina Noack, Hendrik Patzlaff, Felix Naumann, Hazar Harmouch
Inf. Syst.1
2023 A statistical method for predicting quantitative variables in association rule mining
abstract
Association rules encode common patterns and structures identified in datasets. They can be derived by association rule mining (ARM) algorithms. The association rules are human-readable and allow comprehensible predictions, unlike many other types of prediction algorithms. Classical ARM algorithms, like Apriori or FP-growth, cannot process interval or ratio scaled data (quantitative variables) which limits their applicability. We address this restriction in classical ARM algorithms, making it possible to process quantitative variables on the right side of a rule. Our approach is based on applying the Kullback–Leibler divergence (KLD) to identify a rule which holistically considers complete data distributions instead of using only summary statistics. We demonstrate the new approach by using, among others, the example of predicting the length of stay of intensive care patients. The length of stay describes the number of days a patient spends in the intensive care unit. In addition, we further demonstrate our approach by predicting the credit score of bank customers and the contract duration of customers of a fictional telco company based on two publically available datasets. This paper shows a new approach for predicting quantitative variables in ARM. We demonstrate the new approach using the FP-growth algorithm.
Sedir Mohammed, Kerstin Rubarth, Sophie K. Piper, Fridtjof Schiefenhövel, Johann-Christoph Freytag, Felix Balzer, Sebastian Boie
Inf. Syst.1