EDBT 2026 Demo / reviewers in the wild / expert
Soror Sahri
dblp:79/5731
· DBLP profile ↗
10ranked-venue papers in the field
2as first author
7since 2021 · last 2026
0000-0002-1554-7565ORCID · corroborated
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 6 (2 first)Information Retrieval & Web Search · 2Big Data, Cloud & Distributed Data Systems · 1Knowledge Engineering, Semantic Web & Information Systems · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OMNIA: Closing the Loop by Leveraging LLMs for Knowledge Graphs Completion
Frédéric Ieng, Soror Sahri, Mourad Ouzzani, Massinissa Hammaz, Salima Benbernou, Hanieh Khorashadizadeh, Sven Groppe, Farah Benamara |
ICDE | 2 |
| 2025 | Utilizing Quantum Computing to Improve the Quality of Data
Valter Uotila, Soror Sahri, Sven Groppe |
ADBIS | 2 |
| 2025 | Automated Data Quality Validation in an End-to-End GNN Framework
Sijie Dong, Soror Sahri, Themis Palpanas, Qitong Wang 0003 |
EDBT | 2 |
| 2025 | EcoRAG: A Multi-hop Economic QA Benchmark for Retrieval Augmented Generation Using Knowledge Graphs
Hanieh Khorashadizadeh, Sanju Mishra, Farah Benamara, Nandana Mihindukulasooriya, Jinghua Groppe, Soror Sahri, Morteza Kamaladdini Ezzabady, Frédéric Ieng, Sven Groppe |
NLDB (2) | 6 |
| 2024 | Towards Generating High-Quality Knowledge Graphs by Leveraging Large Language Models
Morteza Kamaladdini Ezzabady, Frédéric Ieng, Hanieh Khorashadizadeh, Farah Benamara, Sven Groppe, Soror Sahri |
NLDB (1) | 6 |
| 2024 | Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesabstractDespite the increasing success of Machine Learning (ML) techniques in real-world applications, their maintenance over time remains challenging. In particular, the prediction accuracy of deployed ML models can suffer due to significant changes between training and serving data over time, known as data drift. Traditional data drift solutions primarily focus on detecting drift, and then retraining the ML models, but do not discern whether the detected drift is harmful to model performance. In this paper, we observe that not all data drifts lead to degradation in prediction accuracy. We then introduce a novel approach for identifying portions of data distributions in serving data where drift can be potentially harmful to model performance, which we term Data Distributions with Low Accuracy (DDLA). Our approach, using decision trees, precisely pinpoints low-accuracy zones within ML models, especially Blackbox models. By focusing on these DDLAs, we effectively assess the impact of data drift on model performance and make informed decisions in the ML pipeline. In contrast to existing data drift techniques, we advocate for model retraining only in cases of harmful drifts that detrimentally affect model performance. Through extensive experimental evaluations on various datasets and models, our findings demonstrate that our approach significantly improves cost-efficiency over baselines, while achieving comparable accuracy. Sijie Dong, Qitong Wang 0003, Soror Sahri, Themis Palpanas, Divesh Srivastava |
Proc. VLDB Endow. | 3 |
| 2021 | Customized Eager-Lazy Data Cleansing for Satisfactory Big Data VeracityabstractBig data systems are becoming mainstream for big data management either for batch processing or real-time processing. In order to extract insights from data, quality issues are very important to address, particularly. A veracity assessment model is consequently needed. In this paper, we propose a model which ties quality of datasets and quality of query resultsets. We particularly examine quality issues raised by a given dataset, order attributes along their fitness for use and correlate veracity metrics to business queries. We validate our work using the open dataset NYC taxi’ trips. Soror Sahri, Rim Moussa |
IDEAS | 1 |
| 2015 | CrowdMD: Crowdsourcing-based approach for deduplicationabstractMatching dependencies (MDs) were recently introduced as quality rules for data cleaning and entity resolution. They are rules that specify what values should be considered duplicates, and have to be matched. Defining such quality rules on a database instance, is a very expensive and a time consuming process, and requires huge efforts to analyse the whole database. In this demo paper, we present CrowdMD, a hybrid machine-crowd system for generating MDs. It first asks the crowd to determine whether a given pair, from training sample pairs, match or not. Then, it uses data mining techniques to generate attributes constituting an MD. Using a Restaurant database, we will show how the crowders can help to generate MDs by labelling the training sample through the CrowdMD user interface and how MDs can be mined from this training set. Asma Abboura, Soror Sahri, Mourad Ouziri, Salima Benbernou |
IEEE BigData | 2 |
| 2014 | Summary-Based Pattern Tableaux Generation for Conditional Functional Dependencies in Distributed Data
Soror Sahri, Mourad Ouziri, Salima Benbernou |
DEXA (1) | 1 |
| 2012 | Timed Privacy-Aware Business ProtocolsabstractWeb services privacy issues have been attracting more and more attention in the past years. Since the number of Web services-based business applications is increasing, the demands for privacy enhancing technologies for Web services will also be increasing in the future. In this paper, we investigate an extension of business protocols, i.e. the specification of which message exchange sequences are supported by the web service, in order to accommodate privacy aspects and time-related properties. For this purpose we introduce the notion of Timed Privacy-aware Business Protocols (TPBPs). We also discuss TPBP properties can be checked and we describe their verification process. Karima Mokhtari-Aslaoui, Salima Benbernou, Soror Sahri, Vasilios Andrikopoulos, Frank Leymann, Mohand-Said Hacid |
Int. J. Cooperative Inf. Syst. | 3 |