EDBT 2026 Demo / reviewers in the wild / expert
Sherif Sakr
dblp:s/SherifSakr
· DBLP profile ↗
62ranked-venue papers
16as first author
7since 2021 · last 2025
0000-0002-2503-523XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 34 · 9 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 3 since 2021Systems, architecture and hardware · 7 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 1 first-authorComputer networks · 2Human-computer interaction and ubiquitous computing · 2Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | To tune or not to tune? An approach for recommending important hyperparameters for classification and clustering algorithms
Radwa El Shawi, Mohamadjavad Bahmani, Sherif Sakr |
Future Gener. Comput. Syst. | 3 |
| 2024 | AutoMLBench: A comprehensive experimental evaluation of automated machine learning frameworksabstractWith the booming demand for machine learning applications, it has been recognized that the number of knowledgeable data scientists can not scale with the growing data volumes and application needs in our digital world. In response to this demand, several automated machine learning (AutoML) frameworks have been developed to fill the gap of human expertise by automating the process of building machine learning pipelines. Each framework comes with different heuristics-based design decisions. In this study, we present a comprehensive evaluation and comparison of the performance characteristics of six popular AutoML frameworks, namely, AutoWeka, AutoSKlearn, TPOT, Recipe, ATM and SmartML across 100 data sets from established AutoML benchmark suites. Our experimental evaluation considers different aspects for its comparison, including the performance impact of several design decisions, including time budget, size of search space, meta-learning, and ensemble construction. The results of our study reveal various interesting insights that can significantly guide and impact the design of AutoML frameworks. Hassan Eldeeb, Mohamed Maher 0001, Radwa El Shawi, Sherif Sakr |
Expert Syst. Appl. | 4 |
| 2022 | D2IA: User-defined interval analytics on distributed streams
Ahmed Awad 0001, Riccardo Tommasini 0001, Samuele Langhi, Mahmoud Kamel, Emanuele Della Valle, Sherif Sakr |
Inf. Syst. | 6 |
| 2021 | cSmartML: A Meta Learning-Based Framework for Automated Selection and Hyperparameter Tuning for ClusteringabstractNovel technologies in automated machine learning ease the complexity of algorithm selection and hyper-parameter optimization. However, these are usually restricted to supervised learning tasks such as classification and regression, while unsupervised learning remains a largely unexplored problem. In this paper, we offer a solution for automating machine learning specifically for the case of unsupervised learning with clustering, in a domain-agnostic manner. This is achieved through a combination of state-of-the-art processes based on meta-learning for algorithm and evaluation criteria selection, and evolutionary algorithm for hyper-parameter tuning. We introduce a robust and scalable interactive tool, named cSmartML, built on scikit-learn with 8 clustering algorithms. In order to capture more than a single measure of goodness of the output clustering solution, cSmartML optimizes multiple objective functions. A pareto-approach evaluates each objective simultaneously for each clustering solution. On each of the 27 real and synthetic benchmark datasets, we show that the performance of cSmartML is often much better than using standard selection and hyper-parameter optimization methods. In addition, experimentation reveals that cSmartML takes advantage of the defined objective functions on multi-objective functions framework. Radwa El Shawi, Hudson Lekunze, Sherif Sakr |
IEEE BigData | 3 |
| 2021 | Towards Automated Concept-based Decision TreeExplanations for CNNs
Radwa El Shawi, Youssef Sherif, Sherif Sakr |
EDBT | 3 |
| 2021 | Interpretability in healthcare: A comparative study of local machine learning interpretability techniquesabstractAbstract Although complex machine learning models (eg, random forest, neural networks) are commonly outperforming the traditional and simple interpretable models (eg, linear regression, decision tree), in the healthcare domain, clinicians find it hard to understand and trust these complex models due to the lack of intuition and explanation of their predictions. With the new general data protection regulation (GDPR), the importance for plausibility and verifiability of the predictions made by machine learning models has become essential. Hence, interpretability techniques for machine learning models are an area focus of research. In general, the main aim of these interpretability techniques is to shed light and provide insights into the prediction process of the machine learning models and to be able to explain how the results from the prediction was generated. A major problem in this context is that both the quality of the interpretability techniques and trust of the machine learning model predictions are challenging to measure. In this article, we propose four fundamental quantitative measures for assessing the quality of interpretability techniques— similarity , bias detection , execution time , and trust . We present a comprehensive experimental evaluation of six recent and popular local model agnostic interpretability techniques, namely, LIME , SHAP , Anchors , LORE , ILIME “ and MAPLE on different types of real‐world healthcare data. Building on previous work, our experimental evaluation covers different aspects for its comparison including identity , stability , separability , similarity , execution time , bias detection , and trust . The results of our experiments show that MAPLE achieves the highest performance for the identity across all data sets included in this study, while LIME achieves the lowest performance for the identity metric. LIME achieves the highest performance for the separability metric across all data sets. On average, SHAP has the smallest average time to output explanation across all data sets included in this study. For detecting the bias, SHAP and MAPLE enable the participants to better detect the bias. For the trust metric, Anchors achieves the highest performance on all data sets included in this work. Radwa El Shawi, Youssef Sherif, Mouaz H. Al-Mallah, Sherif Sakr |
Comput. Intell. | 4 |
| 2021 | SDDM: an interpretable statistical concept drift detection method for data streams
Simona Micevska, Ahmed Awad 0001, Sherif Sakr |
J. Intell. Inf. Syst. | 3 |
| 2020 | Declarative Languages for Big Streaming Data
Riccardo Tommasini 0001, Sherif Sakr, Emanuele Della Valle, Hojjat Jafarpour |
EDBT | 2 |
| 2020 | DISGD: A Distributed Shared-nothing Matrix Factorization for Large Scale Online Recommender Systems
Heidy Hazem, Ahmed Awad 0001, Ahmed Hassan Yousef, Sherif Sakr |
EDBT | 4 |
| 2020 | D-SmartML: A Distributed Automated Machine Learning FrameworkabstractNowadays, machine learning is playing a crucial role in harnessing the value of massive data amount currently produced every day. The process of building a high-quality machine learning model is an iterative, complex and time-consuming process that requires solid knowledge about the various machine learning algorithms in addition to having a good experience with effectively tuning their hyper-parameters. With the booming demand for machine learning applications, it has been recognized that the number of knowledgeable data scientists can not scale with the growing data volumes and application needs in our digital world. Therefore, recently, several automated machine learning (AutoML) frameworks have been developed by automating the process of Combined Algorithm Selection and Hyper-parameter tuning (CASH). However, a main limitation of these frameworks is that they have been built on top of centralized machine learning libraries (e.g. scikit-learn) that can only work on a single node and thus they are not scalable to process and handle large data volumes. To tackle this challenge, we demonstrate D-SmartML, a distributed AutoML framework on top of Apache Spark, a distributed data processing framework. Our framework is equipped with a meta learning mechanism for automated algorithm selection and supports three different automated hyper-parameter tuning techniques: distributed grid search, distributed random search and distributed hyperband optimization. We will demonstrate the scalability of our framework on handling large datasets. In addition, we will show how our framework outperforms the-state-of-the-art framework for distributed AutoML optimization, TransmogrifAI. Ahmed Abd Elrahman, Mohamed ElHelw, Radwa El Shawi, Sherif Sakr |
ICDCS | 4 |
| 2020 | Process Mining over Unordered Event StreamsabstractProcess mining is no longer limited to the one-off analysis of static event logs extracted from a single enterprise system. Rather, process mining may strive for immediate insights based on streams of events that are continuously generated by diverse information systems. This requires online algorithms that, instead of keeping the whole history of event data, work incrementally and update analysis results upon the arrival of new events. While such online algorithms have been proposed for several process mining tasks, from discovery through conformance checking to time prediction, they all assume that an event stream is ordered, meaning that the order of event generation coincides with their arrival at the analysis engine. Yet, once events are emitted by independent, distributed systems, this assumption may not hold true, which compromises analysis accuracy. In this paper, we provide the first contribution towards handling unordered event streams in process mining. Specifically, we formalize the notion of out-of-order arrival of events, where an online analysis algorithm needs to process events in an order different from their generation. Using directly-follows graphs as a basic model for many process mining tasks, we provide two approaches to handle such unorderedness, either through buffering or speculative processing. Our experiments with synthetic and real-life event data show that these techniques help mitigate the accuracy loss induced by unordered streams. Ahmed Awad 0001, Matthias Weidlich 0001, Sherif Sakr |
ICPM | 3 |
| 2020 | On Teaching Web Stream Processing - Lessons Learned
Riccardo Tommasini 0001, Emanuele Della Valle, Marco Balduini, Sherif Sakr |
ICWE | 4 |
| 2020 | A First Step Towards a Streaming Linked Data Life-Cycle
Riccardo Tommasini 0001, Mohamed Ragab 0001, Alessandro Falcetta, Emanuele Della Valle, Sherif Sakr |
ISWC (2) | 5 |
| 2020 | Benchmarking big data systems: A survey
Fuad Bajaber, Sherif Sakr, Omar Batarfi, Abdulrahman H. Altalhi, Ahmed Barnawi |
Comput. Commun. | 2 |
| 2020 | The views, measurements and challenges of elasticity in the cloud: A review
Ahmed Barnawi, Sherif Sakr, Wenjing Xiao, Abdullah Al-Barakati |
Comput. Commun. | 2 |
| 2019 | ILIME: Local and Global Interpretable Model-Agnostic Explainer of Black-Box Decision
Radwa El Shawi, Youssef Sherif, Mouaz H. Al-Mallah, Sherif Sakr |
ADBIS | 4 |
| 2019 | D ^2 2 IA: Stream Analytics on User-Defined Event Intervals
Ahmed Awad 0001, Riccardo Tommasini 0001, Mahmoud Kamel, Emanuele Della Valle, Sherif Sakr |
CAiSE | 5 |
| 2019 | LDLCT An Instance-Based Framework for Lesion Detection on Lung CT ScansabstractMedical images have played a crucial role in transforming diagnostic medicine by providing the medical staff with several insights into the health status of every patient. Diagnosis of medical images is a very subjective process which is solely based on the physicians' expertise. In particular, lesion detection is a challenging task due to the various types, shapes and sizes of the lesions in the different organs. Thus, there is a crucial need to build computer-aided frameworks for automated analysis of medical images. In this paper, we present LDLCT, an instance-based framework for automated Lesion Detection on Lung CT Scans. The framework employs Restricted Botlzman Machines (RBM) network for feature extraction stage as an unsupervised feature mapping. The Random Forest (RF) classifier is used to distinguish between pixels from lesion and normal regions. Finally, a post-processing stage is implemented to filter out the false positive candidate lesions. In this study, we select 909 slices with 917 lesions from DeepLesion data set. LDLCT achieves lesion detection sensitivity of 89% with 5 false positives per image, the majority of them can be easily detected by the medical staff. Tarun Khajuria, Eman Badr, Mouaz H. Al-Mallah, Sherif Sakr |
CBMS | 4 |
| 2019 | Interpretability in HealthCare A Comparative Study of Local Machine Learning Interpretability TechniquesabstractAlthough complex machine learning models (e.g., Random Forest, Neural Networks) are commonly outperforming the traditional simple interpretable models (e.g., Linear Regression, Decision Tree), in the healthcare domain, clinicians find it hard to understand and trust these complex models due to the lack of intuition and explanation of their predictions. With the new General Data Protection Regulation (GDPR), the importance for plausibility and verifiability of the predictions made by machine learning models has become essential. To tackle this challenge, recently, several machine learning interpretability techniques have been developed and introduced. In general, the main aim of these interpretability techniques is to shed light and provide insights into the predictions process of the machine learning models and explain how the model predictions have resulted. However, in practice, assessing the quality of the explanations provided by the various interpretability techniques is still questionable. In this paper, we present a comprehensive experimental evaluation of three recent and popular local model agnostic interpretability techniques, namely, LIME, SHAP and Anchors on different types of real-world healthcare data. Our experimental evaluation covers different aspects for its comparison including identity, stability, separability, similarity, execution time and bias detection. The results of our experiments show that LIME achieves the lowest performance for the identity metric and the highest performance for the separability metric across all datasets included in this study. On average, SHAP has the smallest average time to output explanation across all datasets included in this study. For detecting the bias, SHAP enables the participants to better detect the bias. Radwa El Shawi, Youssef Sherif, Mouaz H. Al-Mallah, Sherif Sakr |
CBMS | 4 |
| 2019 | Adaptive Watermarks: A Concept Drift-based Approach for Predicting Event-Time Progress in Data Streams
Ahmed Awad 0001, Jonas Traub, Sherif Sakr |
EDBT | 3 |
| 2019 | SmartML: A Meta Learning-Based Framework for Automated Selection and Hyperparameter Tuning for Machine Learning AlgorithmsabstractInternational audience Mohamed Maher 0001, Sherif Sakr |
EDBT | 2 |
| 2019 | MINARET: A Recommendation Framework for Scientific ReviewersabstractInternational audience Sherif Sakr, Mohamed Ragab 0001, Mohamed Maher 0001, Ahmed Awad 0001 |
EDBT | 1 |
| 2019 | Editorial for Special issue of FGCS special issue on "Benchmarking big data systems"
Sherif Sakr, Albert Y. Zomaya, Athanasios V. Vasilakos |
Future Gener. Comput. Syst. | 1 |
| 2018 | HDM-MC in-Action: A Framework for Big Data Analytics across Multiple ClustersabstractBig data are increasingly collected and stored in a highly distributed infrastructures due to the development of several emerging technologies including sensor network, cloud computing, IoT and mobile computing among many other emerging technologies. In practice, the majority of existing big data processing frameworks (e.g., Hadoop, Spark, Flink) are designed based on the single-cluster setup with the assumptions of centralized management and homogeneous connectivity which makes them sub-optimal and sometimes infeasible to be applied for scenarios that require implementing data analytics jobs on highly distributed data sets (across racks, data centers or multi organizations). We demonstrate HDM-MC, a big data processing framework that is designed to enable the capability of performing large scale data analytics across multi-clusters with minimum extra overhead due to additional scheduling requirements. We describe the architecture and realization of the system using a step-by-step example scenario. Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Sung Une Lee, Huijun Wu 0001 |
ICDCS | 2 |
| 2018 | HDM: A Composable Framework for Big Data ProcessingabstractOver the past years, frameworks such as MapReduce and Spark have been introduced to ease the task of developing big data programs and applications. However, the jobs in these frameworks are roughly defined and packaged as executable jars without any functionality being exposed or described. This means that deployed jobs are not natively composable and reusable for subsequent development. Besides, it also hampers the ability for applying optimizations on the data flow of job sequences and pipelines. In this paper, we present the Hierarchically Distributed Data Matrix (HDM) which is a functional, strongly-typed data representation for writing composable big data applications. Along with HDM, a runtime framework is provided to support the execution, integration and management of HDM applications on distributed infrastructures. Based on the functional data dependency graph of HDM, multiple optimizations are applied to improve the performance of executing HDM jobs. The experimental results show that our optimizations can achieve improvements between 10 to 40 percent of the Job-Completion-Time for different types of applications when compared with the current state of art, Apache Spark. Dongyao Wu, Liming Zhu 0001, Qinghua Lu 0001, Sherif Sakr |
IEEE Trans. Big Data | 4 |
| 2018 | A Differentiated Caching Mechanism to Enable Primary Storage Deduplication in CloudsabstractExisting primary deduplication techniques either use inline caching to exploit locality in primary workloads or use post-processing deduplication to avoid the negative impact on I/O performance. However, neither of them works well in the cloud servers running multiple services for the following two reasons: First, the temporal locality of duplicate data writes varies among primary storage workloads, which makes it challenging to efficiently allocate the inline cache space and achieve a good deduplication ratio. Second, the post-processing deduplication does not eliminate duplicate I/O operations that write to the same logical block address as it is performed after duplicate blocks have been written. A hybrid deduplication mechanism is promising to deal with these problems. Inline fingerprint caching is essential to achieving efficient hybrid deduplication. In this paper, we present a detailed analysis of the limitations of using existing caching algorithms in primary deduplication in the cloud. We reveal that existing caching algorithms either perform poorly or incur significant memory overhead in fingerprint cache management. To address this, we propose a novel fingerprint caching mechanism that estimates the temporal locality of duplicates in different data streams and prioritizes the cache allocation based on the estimation. We integrate the caching mechanism and build a hybrid deduplication system. Our experimental results show that the proposed mechanism provides significant improvement for both deduplication ratio and overhead reduction. Huijun Wu 0001, Chen Wang 0008, Yinjin Fu, Sherif Sakr, Kai Lu 0001, Liming Zhu 0001 |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2017 | Towards Big Data Analytics across Multiple ClustersabstractBig data are increasingly collected and stored in a highly distributed infrastructures due to the development of sensor network, cloud computing, IoT and mobile computing among many other emerging technologies. In practice, the majority of existing big-data-processing frameworks (e.g., Hadoop and Spark) are designed based on the single-cluster setup with the assumptions of centralized management and homogeneous connectivity which makes them sub-optimal and sometimes infeasible to apply for scenarios that require implementing data analytics jobs on highly distributed data sets (across racks, data centers or multi-organizations). In order to tackle this challenge, we present HDM-MC, a multi-cluster big data processing framework which is designed to enable the capability of performing large scale data analytics across multi-clusters with minimum extra overhead due to additional scheduling requirements. In this paper, we present the architecture and realization of the system. In addition, we evaluate the performance of our framework in comparison to other state-of-art single cluster big data processing frameworks. Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Huijun Wu 0001 |
CCGrid | 2 |
| 2017 | HDM: Optimized Big Data Processing with Data Provenance
Dongyao Wu, Sherif Sakr, Liming Zhu 0001 |
EDBT | 2 |
| 2017 | On business process monitoring using cross-flow coordination
Zakaria Maamar, Noura Faci, Mohamed Sellami, Khouloud Boukadi, Fadwa Yahya, Ahmed Barnawi, Sherif Sakr |
Serv. Oriented Comput. Appl. | 7 |
| 2016 | Big Data 2.0 Processing Systems: Taxonomy and Open Challenges
Fuad Bajaber, Radwa El Shawi, Omar Batarfi, Abdulrahman H. Altalhi, Ahmed Barnawi, Sherif Sakr |
J. Grid Comput. | 6 |
| 2016 | Network-based social coordination of business processes
Zakaria Maamar, Noura Faci, Sherif Sakr, Mohamed Boukhebouze, Ahmed Barnawi |
Inf. Syst. | 3 |
| 2015 | Composable and efficient functional big data processing frameworkabstractOver the past years, frameworks such as MapReduce and Spark have been introduced to ease the task of developing big data programs and applications. However, the jobs in these frameworks are roughly defined and packaged as executable jars without any functionality being exposed or described. This means that deployed jobs are not natively composable and reusable for subsequent development. Besides, it also hampers the ability for applying optimizations on the data flow of job sequences and pipelines. In this paper, we present the Hierarchically Distributed Data Matrix (HDM) which is a functional, strongly-typed data representation for writing composable big data applications. Along with HDM, a runtime framework is provided to support the execution of HDM applications on distributed infrastructures. Based on the functional data dependency graph of HDM, multiple optimizations are applied to improve the performance of executing HDM jobs. The experimental results show that our optimizations can achieve improvements of between 10% to 60% of the Job-Completion-Time for different types of operation sequences when compared with the current state of art, Apache Spark. Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Qinghua Lu 0001 |
IEEE BigData | 2 |
| 2015 | Liquid Benchmarking: A Platform for Democratizing the Performance Evaluation ProcessabstractPerformances evaluation, reproducibility and benchmarking represent crucial aspects for assessing the practical impact of research results in the computer science eld. In spite of all the benets (e.g., increasing impact, increasing visibility, improving the research quality) that can be gained from performing extensive experimental evaluation or providing reproducible software artifacts and detailed description of experimental setup, the required eort for achiev Sherif Sakr, Amin Shafaat, Fuad Bajaber, Ahmed Barnawi, Omar Batarfi, Abdulrahman H. Altalhi |
EDBT | 1 |
| 2015 | DREAM: Distributed RDF Engine with Adaptive Query Planner and Minimal CommunicationabstractThe Resource Description Framework (RDF) and SPARQL query language are gaining wide popularity and acceptance. In this paper, we present DREAM, a distributed and adaptive RDF system. As opposed to existing RDF systems, DREAM avoids partitioning RDF datasets and partitions only SPARQL queries. By not partitioning datasets, DREAM offers a general paradigm for different types of pattern matching queries, and entirely averts intermediate data shuffling (only auxiliary data are shuffled). Besides, by partitioning queries, DREAM presents an adaptive scheme, which automatically runs queries on various numbers of machines depending on their complexities. Hence, in essence DREAM combines the advantages of the state-of-the-art centralized and distributed RDF systems, whereby data communication is avoided and cluster resources are aggregated. Likewise, it precludes their disadvantages, wherein system resources are limited and communication overhead is typically hindering. DREAM achieves all its goals via employing a novel graph-based, rule-oriented query planner and a new cost model. We implemented DREAM and conducted comprehensive experiments on a private cluster and on the Amazon EC2 platform. Results show that DREAM can significantly outperform three related popular RDF systems. Mohammad Hammoud, Dania Abed Rabbou, Reza Nouri, Amin Beheshti, Sherif Sakr |
Proc. VLDB Endow. | 5 |
| 2015 | A Framework for Consumer-Centric SLA Management of Cloud-Hosted DatabasesabstractService Level Agreements (SLA) represent the contract which captures the agreed upon guarantees between a service provider and its customers. The specifications of existing service level agreements (SLA) for cloud services are not designed to flexibly handle even relatively straightforward performance and technical requirements of consumer applications. In this article, we present a novel approach for SLA-based management of cloud-hosted databases from the consumer perspective. We present an end-to-end framework for consumer-centric SLA management of cloud-hosted databases. The framework facilitates adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. In this framework, the SLA of the consumer applications are declaratively defined in terms of goals which are subjected to a number of constraints that are specific to the application requirements. The framework continuously monitors the application-defined SLA and automatically triggers the execution of necessary corrective actions (scaling out/in the database tier) when required. The framework is database platform-agnostic, uses virtualization-based database replication mechanisms, and requires zero source code changes of the cloud-hosted software applications. The experimental results demonstrate the effectiveness of our SLA-based framework in providing the consumer applications with the required flexibility for achieving their SLA requirements. Liang Zhao 0009, Sherif Sakr, Anna Liu |
IEEE Trans. Serv. Comput. | 2 |
| 2014 | CDPort: A Framework of Data Portability in Cloud PlatformsabstractOne of the main advantages of the cloud computing paradigm is that it simplifies the time-consuming processes of hardware provisioning, hardware procurement and software deployment. Currently, we are witnessing a proliferation in the number of cloud-hosted applications. However, one of the important challenges of the cloud computing paradigm that may hurt the growth of this technology is the interoperability and portability between cloud platforms. The developers and cloud users may lock-in to the first cloud they choose, or face difficulties when they have to move their data or software from one cloud platform to another. In this paper, we focus on the challenge of data portability between different cloud-based data storage services. In particular, we propose a common data model and a standardized API for the new generation of cloud-based NoSQL databases. The initial implementation of our framework covers three of the most popular NoSQL systems, namely, Google Datastore, Amazon SimpleDB and MongoDB. However, our framework is designed in a flexible way that it can be easily extended to support other NoSQL systems. Furthermore, our framework is equipped with tools that support the conversion, transformation and exchange of the data which is stored on the supported NoSQL databases of the framework. Finally, we describe the design and the proof-of-concept implementation of our framework using a case study. Ebtesam Ahmad Alomari, Ahmed Barnawi, Sherif Sakr |
iiWAS | 3 |
| 2014 | Hybrid query execution engine for large attributed graphs
Sherif Sakr, Sameh Elnikety, Yuxiong He |
Inf. Syst. | 1 |
| 2013 | Improving Availability of Cloud-Based Applications through Deployment ChoicesabstractDeployment choices are critical in determining the availability of applications running in a cloud. But choosing good deployment for various software application components into virtual machines is a challenging task because of potential sharing of components among applications and potential interference from multi-tenancy. This paper presents an approach for improving the availability guarantee of software applications by optimizing the availability, performance and monetary cost trade-offs of different deployment choices. Our approach explicitly considers different classes of application requests during the decision process. The results of our experimental evaluation show that the approach can effectively improve the availability guarantees with little or negligible increase in the performance and monetary cost of the deployment choice. Jim Zhanwen Li, Qinghua Lu 0001, Liming Zhu 0001, Leonard J. Bass, Xiwei Xu 0001, Sherif Sakr, Paul L. Bannerman, Anna Liu |
IEEE CLOUD | 6 |
| 2013 | Incorporating Uncertainty into In-Cloud Application Deployment Decisions for AvailabilityabstractCloud consumers have a variety of deployment related techniques, such as auto-scaling policies and recovery strategies, for dealing with the uncertainties in the cloud. Uncertainties can be characterized as stochastic (such as failures, disasters, and workload spikes) and subjective (such as choice among various deployment options). Cloud consumers must consider both stochastic and subjective uncertainties. Analytic support for consumers in selecting appropriate techniques and setting the required parameters in the face of different types of uncertainty is currently limited. In this paper, we propose a set of application availability analysis models that capture subjective uncertainties in addition to stochastic uncertainties. We built and validated the models by using industry best practices on deployment, and actual commercial products for disaster recovery and live migration. Our results show that the models permit more informed and quantitative availability analysis than industry best practices under a wide range of scenarios. Qinghua Lu 0001, Xiwei Xu 0001, Liming Zhu 0001, Leonard J. Bass, Jim Zhanwen Li, Sherif Sakr, Paul L. Bannerman, Anna Liu |
IEEE CLOUD | 6 |
| 2013 | Consumer-centric SLA manager for cloud-hosted databasesabstractWe present an end-to-end framework for consumer-centric SLA management of virtualized database servers. The framework facilitates adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. In this framework, the SLA of the consumer applications are declaratively defined in terms of goals which are subjected to a number of constraints that are specific to the application requirements. The framework continuously monitors the application-defined SLA and automatically triggers the execution of necessary corrective actions (scaling out/in the database tier) when required. The framework is database platform-agnostic, uses virtualization-based database replication mechanisms and requires zero source code changes of the cloud-hosted application. Liang Zhao 0009, Sherif Sakr, Anna Liu |
CIKM | 2 |
| 2013 | Is Your Cloud-Hosted Database Truly Elastic?abstractElasticity has been recognized as one of the most appealing features for users of cloud services. It represents the ability to dynamically and rapidly scale up or down the allocated computing resources on demand. In practice, it is difficult to understand the elasticity requirements of a given application and workload, and to assess if the elasticity provided by a cloud service will meet these requirements. In this experience paper, we take the position that a deep understanding of the capabilities of cloud-hosted database services is a crucial requirement for cloud users in order to bring forward the vision of deploying data-intensive applications on cloud platforms. We argue that it is important that cloud users become able to paint a comprehensive picture of the relationship between the capabilities of the different type of cloud database services, the application characteristics and workloads, and the geographical distribution of the application clients and the underlying database replicas. We discuss the current elasticity capabilities of the different categories of cloud database services and identify some of the main challenges for deploying a truly elastic database tier on cloud environments. Finally, we propose a benchmarking mechanism that can evaluate the elasticity capabilities of cloud database services in different application scenarios and workloads. Sherif Sakr, Anna Liu |
SERVICES | 1 |
| 2013 | Modeling performance of a parallel streaming engine: bridging theory and costsabstractWhile data are growing at a speed never seen before, parallel computing is becoming more and more essential to process this massive volume of data in a timely manner. Therefore, recently, concurrent computations have been receiving increasing attention due to the widespread adoption of multi-core processors and the emerging advancements of cloud computing technology. The ubiquity of mobile devices, location services, and sensor pervasiveness are examples of new scenarios that have created the crucial need for building scalable computing platforms and parallel architectures to process vast amounts of generated streaming data. In practice, efficiently operating these systems is hard due to the intrinsic complexity of these architectures and the lack of a formal and in-depth knowledge of the performance models and the consequent system costs. The Actor Model theory has been presented as a mathematical model of con- current computation that had enormous success in practice and inspired a number of contemporary work in this area. Recently, the Storm system has been presented as a realization of the principles of the Actor Model theory in the context of the large scale processing of streaming data. In this paper, we present, to the best of our knowledge, the first set of models that formalize the performance characteristics of a practical distributed, parallel and fault-tolerant stream processing system that follows the Actor Model theory. In particular, we model the characteristics of the data flow, the data processing and the system management costs at a fine granularity within the different steps of executing a distributed stream processing job. Finally, we present an experimental validation of the described performance models using the Storm system. Ivan Bedini, Sherif Sakr, Bart Theeten, Alessandra Sala, Peter Cogan |
ICPE | 2 |
| 2012 | SLA-Based and Consumer-centric Dynamic Provisioning for Cloud DatabasesabstractOne of the main advantages of the cloud computing paradigm is that it simplifies the time-consuming processes of hardware provisioning, hardware purchasing and software deployment. Currently, we are witnessing a proliferation in the number of cloud-hosted applications with a tremendous increase in the scale of the data generated as well as being consumed by such applications. Cloud-hosted database systems powering these applications form a critical component in the software stack of these applications. Service Level Agreements (SLA) represent the contract which captures the agreed upon guarantees between a service provider and its customers. The specifications of existing service level agreement (SLA) for cloud services are not designed for flexibly handling even relatively straightforward performance and technical requirements of consumer applications. The concerns of consumers for cloud services regarding the SLA management of their hosted applications within the cloud environments will gain increasing importance as cloud computing becomes more pervasive. This paper introduces the notion, challenges and the importance of SLA-based provisioning and cost management for cloud-hosted databases from the consumer perspective. We present an end-to-end framework that acts as a middleware which resides between the consumer applications and the cloud-hosted databases. The aim of the framework is to facilitate adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. The experimental results demonstrate that SLA-based provisioning is more adequate for providing consumer applications the required flexibility in achieving their goals. Sherif Sakr, Anna Liu |
IEEE CLOUD | 1 |
| 2012 | Application-Managed Replication Controller for Cloud-Hosted DatabasesabstractData replication is a well-known strategy to achieve the availability, scalability and performance improvement goals in the data management world. However, the cost of maintaining several database replicas always strongly consistent is very high. The CAP theorem shows that a shared-data system can choose at most two out of three properties: consistency, availability, and tolerance to partitions. In practice, most of the cloud-based data management systems tend to overcome the difficulties of distributed replication by relaxing the consistency guarantees of the system. In particular, they implement various forms of weaker consistency models such as eventual consistency. This solution is accepted by many new Web 2.0 applications (e.g. social networks) which could be more tolerant with a wider window of data staleness (replication delay).However, unfortunately, there are no generic application-independent and consumer-centric mechanisms by which software applications can specify and manage to what extent inconsistencies can be tolerated. We introduce an adaptive framework for database replication at the middleware layer of cloud environments. The framework provides flexible mechanisms to enable software applications of keeping several database replicas (that can be hosted in different data centers) with different levels of service level agreements (SLA) for their data freshness. The experimental evaluation demonstrates the effectiveness of our framework in providing the software applications with the required flexibility to achieve and optimize their requirements in terms of overall system throughput, data freshness and invested monetary cost. Liang Zhao 0009, Sherif Sakr, Anna Liu |
IEEE CLOUD | 2 |
| 2012 | G-SPARQL: a hybrid engine for querying large attributed graphsabstractWe propose a SPARQL-like language, G-SPARQL, for querying attributed graphs. The language expresses types of queries which of large interest for applications which model their data as large graphs such as: pattern matching, reachability and shortest path queries. Each query can combine both of structural predicates and value-based predicates (on the attributes of the graph nodes and edges). We describe an algebraic compilation mechanism for our proposed query language which is extended from the relational algebra and based on the basic construct of building SPARQL queries, the Triple Pattern. We describe a hybrid Memory/Disk representation of large attributed graphs where only the topology of the graph is maintained in memory while the data of the graph is stored in a relational database. The execution engine of our proposed query language splits parts of the query plan to be pushed inside the relational database while the execution of other parts of the query plan are processed using memory-based algorithms, as necessary. Experimental results on real datasets demonstrate the efficiency and the scalability of our approach and show that our approach outperforms native graph databases by several factors. Sherif Sakr, Sameh Elnikety, Yuxiong He |
CIKM | 1 |
| 2012 | Trade-Off Analysis of Elasticity Approaches for Cloud-Based Business Applications
Basem Suleiman, Sherif Sakr, Srikumar Venugopal, Wasim Sadiq |
WISE | 2 |
| 2011 | A Query Language for Analyzing Business Processes Execution
Amin Beheshti, Boualem Benatallah, Hamid R. Motahari Nezhad, Sherif Sakr |
BPM | 4 |
| 2011 | Design by Selection: A Reuse-Based Approach for Business Process Modeling
Ahmed Awad 0001, Sherif Sakr, Matthias Kunze 0001, Mathias Weske |
ER | 2 |
| 2011 | CloudDB AutoAdmin: Towards a Truly Elastic Cloud-Based Data StoreabstractIn this paper, we present the design and the architecture of the CloudDB AutoAdmin system which aims to fill the existing gaps between the provided cloud database services and the requirements of the consumer applications. In particular, it focuses on facilitating the job of the cloud database consumers in implementing database applications as distributed, scalable, and elastic services with a minimum effort on the side of the application developer and a limited footprint in the application code. Sherif Sakr, Liang Zhao 0009, Hiroshi Wada, Anna Liu |
ICWS | 1 |
| 2010 | An efficient features-based processing technique for supergraph queriesabstractGraphs are widely used for modeling complicated data such as social networks, chemical compounds, protein interactions, XML documents and multimedia databases. To be able to effectively understand and utilize any collection of graphs, a graph database that efficiently supports elementary querying mechanisms is crucially required. Supergraph query is an important type of graph queries which has many practical applications. Given a graph database D, the answer set of a supergraph query q is computed by retrieving all graphs in D which are fully contained in q. A primary challenge in computing the answers of graph queries is that pair-wise comparisons of graphs are usually hard problems. For example, subgraph isomorphism is known to be NP-complete. Clearly, the success of any graph database application is directly dependent on the efficiency of the graph indexing and query processing mechanisms. In this paper, we study the problem of using the relational infrastructure to achieve an efficient evaluation of supergraph queries. We rely on an effective and efficient layer of features-based summary structures, called graph features knowledge, to reduce the required number of pair-wise graph comparisons and boost the efficiency of query processing. Finally, we conduct an extensive set of experiments on real and synthetic data sets to demonstrate the efficiency and the scalability of our approach. Sherif Sakr, Ghazi Al-Naymat |
IDEAS | 1 |
| 2010 | Efficient and Adaptable Query Workload-Aware Management for RDF Data
Hooran MahmoudiNasab, Sherif Sakr |
WISE | 2 |
| 2010 | A framework for querying graph-based business process modelsabstractWe present a framework for querying and reusing graph-based business process models. The framework is based on a new visual query language for business processes called BPMN-Q. The language addresses processes definitions and extends the standard BPMN visual notations for modeling business processes for its concrete syntax. BPMN-Q is used to query process models by matching a process model graph to a query graph. Moreover, the reusing framework is enhanced with a semantic query expander component. This component provides the users with the flexibility to get not only the perfectly matched process models to their queries but also the models with high similarity. The query engine of the framework is built on top of traditional RDBMS. A novel decomposition based and selectivity-aware relational processing mechanism is employed to achieve an efficient and scalable performance for graph-based BPMN-Q queries. Sherif Sakr, Ahmed Awad 0001 |
WWW | 1 |
| 2010 | Efficient Relational Techniques for Processing Graph Queries
Sherif Sakr, Ghazi Al-Naymat |
J. Comput. Sci. Technol. | 1 |
| 2009 | GraphREL: A Decomposition-Based and Selectivity-Aware Relational Framework for Processing Sub-graph Queries
Sherif Sakr |
DASFAA | 1 |
| 2009 | FeedRank: A Semantic-Based Management System of Web Feeds
Hooran MahmoudiNasab, Sherif Sakr |
IDEAL | 2 |
| 2009 | XML compression techniques: A survey and comparison
Sherif Sakr |
J. Comput. Syst. Sci. | 1 |
| 2009 | Cardinality-Aware Purely Relational XQuery ProcessorabstractRecently, the use of eXtensible Markup Language (XML) continues to grow in popularity, large repositories of XML documents are going to emerge, and users are likely to pose increasingly more complex queries on these data sets. In 2001 XQuery is decided by the World Wide Web Consortium (W3C) as the standard XML query language. In this article, we describe the design and implementation of an efficient and scalable purely relational XQuery processor which translates expressions of the XQuery language into their equivalent SQL evaluation scripts. The experiments of this article demonstrated the efficiency and scalability of our purely relational approach in comparison to the native XML/XQuery functionality supported by conventional RDBMSs and has shown that our purely relational approach for implementing XQuery processor deserves to be pursued further. Sherif Sakr |
J. Database Manag. | 1 |
| 2008 | XSelMark: A Micro-benchmark for Selectivity Estimation Approaches of XML Queries
Sherif Sakr |
DEXA | 1 |
| 2008 | Improving the Relational Evaluation of XML Queries by Means of Path Summaries
Sherif Sakr |
IDEAL | 1 |
| 2008 | Dependable cardinality forecasts for XQueryabstractThough inevitable for effective cost-based query rewriting, the derivation of meaningful cardinality estimates has remained a notoriously hard problem in the context of XQuery. By basing the estimation on a relational representation of the XQuery syntax, we show how existing cardinality estimation techniques for XPath and proven relational estimation machinery can play together to yield dependable forecasts for arbitrary XQuery (sub)expressions. Our approach benefits from a light-weight form of data flow analysis. Abstract domain identifiers guide our query analyzer through the estimation process and allow for informed decisions even in case of deeply nested XQuery expressions. A variant of projection paths [15] provides a versatile interface into which existing techniques for XPath cardinality estimation can be plugged in seamlessly. We demonstrate an implementation of this interface based on data guides. Experiments show how our approach can equally cope with both, structure-and value-based queries. It is robust with respect to intermediate estimation errors, from which we typically found our implementation to recover gracefully. Jens Teubner, Torsten Grust, Sebastian Maneth, Sherif Sakr |
Proc. VLDB Endow. | 4 |
| 2007 | A SQL: 1999 code generator for the pathfinder xquery compilerabstractThe Pathfinder XQuery compiler has been enhanced by a new code generator that can target any SQL:1999-compliant relational database system(RDBMS). This code generator marks an important next step towards truly relational XQuery processing, a branch of database technology that aims to turn RDBMSs into highly efficient XML and XQuery processors without the need to invade the relational database kernel. Pathfinder, a retargetable front-end compiler, translates input XQuery expressions into DAG-shaped relational algebra plans. The code generator then turns these plans into sequences of either SQL:1999 statements or view definitions which jointly implement the (sometimes intricate) XQuery semantics. In a sense, this demonstration thus lets relational algebra and SQL swap their traditional roles in database query processing. The result is a code generator that (1) supports an almost complete dialect of XQuery, (2) can target any RDBMS with a SQL:1999 language interface, and (3) exhibits quite promising performance characteristics when run against high-volume XML data as well as complex XQuery expressions. Torsten Grust, Manuel Mayr, Jan Rittinger, Sherif Sakr, Jens Teubner |
SIGMOD Conference | 4 |
| 2004 | XQuery on SQL Hosts
Torsten Grust, Sherif Sakr, Jens Teubner |
VLDB | 2 |