VLDB 2026 Research / reviewers in the wild / expert
Soror Sahri
dblp:79/5731
· DBLP profile ↗
14ranked-venue papers
3as first author
8since 2021 · last 2026
0000-0002-1554-7565ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 10 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 5 · 2 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | OMNIA: Closing the Loop by Leveraging LLMs for Knowledge Graphs Completion
Frédéric Ieng, Soror Sahri, Mourad Ouzzani, Massinissa Hammaz, Salima Benbernou, Hanieh Khorashadizadeh, Sven Groppe, Farah Benamara |
ICDE | 2 |
| 2025 | Utilizing Quantum Computing to Improve the Quality of Data
Valter Uotila, Soror Sahri, Sven Groppe |
ADBIS | 2 |
| 2025 | Apriori Meets LLMs: Interpretable Rule Mining from Continuous Healthcare DataabstractMining meaningful patterns from numerical healthcare data is challenging, as continuous lab values are difficult to analyze directly and traditional association rule mining often generates arbitrary thresholds. We introduce Threshold-Aware Association Rules (TAAR), a framework that converts continuous lab values into semantic intervals and extracts interpretable rules using an enhanced Apriori algorithm. Large Language Models (LLMs) are employed to refine support and confidence thresholds, filter implausible rules, and produce natural-language explanations. Applied to blood test data, TAAR improves clinical usability, guides actionable follow-up recommendations, and supports informed decision-making. Asma Abboura, Imane Hocine, Yacine Hakimi, Soror Sahri, Grégoire Danoy |
BIBM | 4 |
| 2025 | Automated Data Quality Validation in an End-to-End GNN Framework
Sijie Dong, Soror Sahri, Themis Palpanas, Qitong Wang 0003 |
EDBT | 2 |
| 2025 | EcoRAG: A Multi-hop Economic QA Benchmark for Retrieval Augmented Generation Using Knowledge Graphs
Hanieh Khorashadizadeh, Sanju Mishra, Farah Benamara, Nandana Mihindukulasooriya, Jinghua Groppe, Soror Sahri, Morteza Kamaladdini Ezzabady, Frédéric Ieng, Sven Groppe |
NLDB (2) | 6 |
| 2024 | Towards Generating High-Quality Knowledge Graphs by Leveraging Large Language Models
Morteza Kamaladdini Ezzabady, Frédéric Ieng, Hanieh Khorashadizadeh, Farah Benamara, Sven Groppe, Soror Sahri |
NLDB (1) | 6 |
| 2024 | Efficiently Mitigating the Impact of Data Drift on Machine Learning PipelinesabstractDespite the increasing success of Machine Learning (ML) techniques in real-world applications, their maintenance over time remains challenging. In particular, the prediction accuracy of deployed ML models can suffer due to significant changes between training and serving data over time, known as data drift. Traditional data drift solutions primarily focus on detecting drift, and then retraining the ML models, but do not discern whether the detected drift is harmful to model performance. In this paper, we observe that not all data drifts lead to degradation in prediction accuracy. We then introduce a novel approach for identifying portions of data distributions in serving data where drift can be potentially harmful to model performance, which we term Data Distributions with Low Accuracy (DDLA). Our approach, using decision trees, precisely pinpoints low-accuracy zones within ML models, especially Blackbox models. By focusing on these DDLAs, we effectively assess the impact of data drift on model performance and make informed decisions in the ML pipeline. In contrast to existing data drift techniques, we advocate for model retraining only in cases of harmful drifts that detrimentally affect model performance. Through extensive experimental evaluations on various datasets and models, our findings demonstrate that our approach significantly improves cost-efficiency over baselines, while achieving comparable accuracy. Sijie Dong, Qitong Wang 0003, Soror Sahri, Themis Palpanas, Divesh Srivastava |
Proc. VLDB Endow. | 3 |
| 2021 | Customized Eager-Lazy Data Cleansing for Satisfactory Big Data VeracityabstractBig data systems are becoming mainstream for big data management either for batch processing or real-time processing. In order to extract insights from data, quality issues are very important to address, particularly. A veracity assessment model is consequently needed. In this paper, we propose a model which ties quality of datasets and quality of query resultsets. We particularly examine quality issues raised by a given dataset, order attributes along their fitness for use and correlate veracity metrics to business queries. We validate our work using the open dataset NYC taxi’ trips. Soror Sahri, Rim Moussa |
IDEAS | 1 |
| 2016 | Quality-Based Online Data ReconciliationabstractOne of the main challenges in data matching and data cleaning, in highly integrated systems, is duplicates detection . While the literature abounds of approaches detecting duplicates corresponding to the same real-world entity, most of these approaches tend to eliminate duplicates (wrong information) from the sources, hence leading to what is called data repair. In this article, we propose a framework that automatically detects duplicates at query time and effectively identifies the consistent version of the data, while keeping inconsistent data in the sources. Our framework uses matching dependencies (MDs) to detect duplicates through the concept of data reconciliation rules (DRR) and conditional function dependencies (CFDs) to assess the quality of different attribute values. We also build a duplicate reconciliation index ( DRI ), based on clusters of duplicates detected by a set of DRRs to speed up the online data reconciliation process. Our experiments of a real-world data collection show the efficiency and effectiveness of our framework. Asma Abboura, Soror Sahri, Latifa Baba-hamed, Mourad Ouziri, Salima Benbernou |
ACM Trans. Internet Techn. | 2 |
| 2015 | CrowdMD: Crowdsourcing-based approach for deduplicationabstractMatching dependencies (MDs) were recently introduced as quality rules for data cleaning and entity resolution. They are rules that specify what values should be considered duplicates, and have to be matched. Defining such quality rules on a database instance, is a very expensive and a time consuming process, and requires huge efforts to analyse the whole database. In this demo paper, we present CrowdMD, a hybrid machine-crowd system for generating MDs. It first asks the crowd to determine whether a given pair, from training sample pairs, match or not. Then, it uses data mining techniques to generate attributes constituting an MD. Using a Restaurant database, we will show how the crowders can help to generate MDs by labelling the training sample through the CrowdMD user interface and how MDs can be mined from this training set. Asma Abboura, Soror Sahri, Mourad Ouziri, Salima Benbernou |
IEEE BigData | 2 |
| 2014 | Summary-Based Pattern Tableaux Generation for Conditional Functional Dependencies in Distributed Data
Soror Sahri, Mourad Ouziri, Salima Benbernou |
DEXA (1) | 1 |
| 2014 | DBaaS-Expert: A Recommender for the Selection of the Right Cloud Database
Soror Sahri, Rim Moussa, Darrell D. E. Long, Salima Benbernou |
ISMIS | 1 |
| 2014 | Be a Collaborator and a Competitor in Crowdsourcing SystemabstractCrowd sourcing is emerging as a powerful paradigm to solve a wide range of tedious and complex problems in various enterprise applications. It spawns the issue of finding the unknown collaborative and competitive group of solvers. The formation of collaborative team should provide the best solution and treat that solution as a trade secret avoiding data leak between competitive teams due to reward behind the outsourcing of the issue. The formation of effective competitive teams not only requires adequate skills between members of each team, but also good team connectivity through social network and to provide the best solution and treat that solution as a trade secret avoiding data leak between teams due to reward behind the outsourcing of the issue. In this paper, we propose a data leak aware crowd sourcing system called Social Crowd. We introduce a clustering algorithm that uses social relationships between crowd workers to discover all possible teams while avoiding inter-team data leakage. Iheb Ben Amor, Mourad Ouziri, Soror Sahri, Naouel Karam |
MASCOTS | 3 |
| 2012 | Timed Privacy-Aware Business ProtocolsabstractWeb services privacy issues have been attracting more and more attention in the past years. Since the number of Web services-based business applications is increasing, the demands for privacy enhancing technologies for Web services will also be increasing in the future. In this paper, we investigate an extension of business protocols, i.e. the specification of which message exchange sequences are supported by the web service, in order to accommodate privacy aspects and time-related properties. For this purpose we introduce the notion of Timed Privacy-aware Business Protocols (TPBPs). We also discuss TPBP properties can be checked and we describe their verification process. Karima Mokhtari-Aslaoui, Salima Benbernou, Soror Sahri, Vasilios Andrikopoulos, Frank Leymann, Mohand-Said Hacid |
Int. J. Cooperative Inf. Syst. | 3 |