Sherif Sakr

dblp:s/SherifSakr · DBLP profile ↗
← Back
62ranked-venue papers
16as first author
7since 2021 · last 2025
0000-0002-2503-523XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 34 · 9 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 4 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 3 since 2021Systems, architecture and hardware · 7 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 7 · 1 first-authorComputer networks · 2Human-computer interaction and ubiquitous computing · 2Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2025 To tune or not to tune? An approach for recommending important hyperparameters for classification and clustering algorithms
Radwa El Shawi, Mohamadjavad Bahmani, Sherif Sakr
Future Gener. Comput. Syst.3
2024 AutoMLBench: A comprehensive experimental evaluation of automated machine learning frameworks
abstract
With the booming demand for machine learning applications, it has been recognized that the number of knowledgeable data scientists can not scale with the growing data volumes and application needs in our digital world. In response to this demand, several automated machine learning (AutoML) frameworks have been developed to fill the gap of human expertise by automating the process of building machine learning pipelines. Each framework comes with different heuristics-based design decisions. In this study, we present a comprehensive evaluation and comparison of the performance characteristics of six popular AutoML frameworks, namely, AutoWeka, AutoSKlearn, TPOT, Recipe, ATM and SmartML across 100 data sets from established AutoML benchmark suites. Our experimental evaluation considers different aspects for its comparison, including the performance impact of several design decisions, including time budget, size of search space, meta-learning, and ensemble construction. The results of our study reveal various interesting insights that can significantly guide and impact the design of AutoML frameworks.
Hassan Eldeeb, Mohamed Maher 0001, Radwa El Shawi, Sherif Sakr
Expert Syst. Appl.4
2022 D2IA: User-defined interval analytics on distributed streams
Ahmed Awad 0001, Riccardo Tommasini 0001, Samuele Langhi, Mahmoud Kamel, Emanuele Della Valle, Sherif Sakr
Inf. Syst.6
2021 cSmartML: A Meta Learning-Based Framework for Automated Selection and Hyperparameter Tuning for Clustering
abstract
Novel technologies in automated machine learning ease the complexity of algorithm selection and hyper-parameter optimization. However, these are usually restricted to supervised learning tasks such as classification and regression, while unsupervised learning remains a largely unexplored problem. In this paper, we offer a solution for automating machine learning specifically for the case of unsupervised learning with clustering, in a domain-agnostic manner. This is achieved through a combination of state-of-the-art processes based on meta-learning for algorithm and evaluation criteria selection, and evolutionary algorithm for hyper-parameter tuning. We introduce a robust and scalable interactive tool, named cSmartML, built on scikit-learn with 8 clustering algorithms. In order to capture more than a single measure of goodness of the output clustering solution, cSmartML optimizes multiple objective functions. A pareto-approach evaluates each objective simultaneously for each clustering solution. On each of the 27 real and synthetic benchmark datasets, we show that the performance of cSmartML is often much better than using standard selection and hyper-parameter optimization methods. In addition, experimentation reveals that cSmartML takes advantage of the defined objective functions on multi-objective functions framework.
Radwa El Shawi, Hudson Lekunze, Sherif Sakr
IEEE BigData3
2021 Towards Automated Concept-based Decision TreeExplanations for CNNs
Radwa El Shawi, Youssef Sherif, Sherif Sakr
EDBT3
2021 Interpretability in healthcare: A comparative study of local machine learning interpretability techniques
abstract
Abstract Although complex machine learning models (eg, random forest, neural networks) are commonly outperforming the traditional and simple interpretable models (eg, linear regression, decision tree), in the healthcare domain, clinicians find it hard to understand and trust these complex models due to the lack of intuition and explanation of their predictions. With the new general data protection regulation (GDPR), the importance for plausibility and verifiability of the predictions made by machine learning models has become essential. Hence, interpretability techniques for machine learning models are an area focus of research. In general, the main aim of these interpretability techniques is to shed light and provide insights into the prediction process of the machine learning models and to be able to explain how the results from the prediction was generated. A major problem in this context is that both the quality of the interpretability techniques and trust of the machine learning model predictions are challenging to measure. In this article, we propose four fundamental quantitative measures for assessing the quality of interpretability techniques— similarity , bias detection , execution time , and trust . We present a comprehensive experimental evaluation of six recent and popular local model agnostic interpretability techniques, namely, LIME , SHAP , Anchors , LORE , ILIME “ and MAPLE on different types of real‐world healthcare data. Building on previous work, our experimental evaluation covers different aspects for its comparison including identity , stability , separability , similarity , execution time , bias detection , and trust . The results of our experiments show that MAPLE achieves the highest performance for the identity across all data sets included in this study, while LIME achieves the lowest performance for the identity metric. LIME achieves the highest performance for the separability metric across all data sets. On average, SHAP has the smallest average time to output explanation across all data sets included in this study. For detecting the bias, SHAP and MAPLE enable the participants to better detect the bias. For the trust metric, Anchors achieves the highest performance on all data sets included in this work.
Radwa El Shawi, Youssef Sherif, Mouaz H. Al-Mallah, Sherif Sakr
Comput. Intell.4
2021 SDDM: an interpretable statistical concept drift detection method for data streams
Simona Micevska, Ahmed Awad 0001, Sherif Sakr
J. Intell. Inf. Syst.3
2020 Declarative Languages for Big Streaming Data
Riccardo Tommasini 0001, Sherif Sakr, Emanuele Della Valle, Hojjat Jafarpour
EDBT2
2020 DISGD: A Distributed Shared-nothing Matrix Factorization for Large Scale Online Recommender Systems
Heidy Hazem, Ahmed Awad 0001, Ahmed Hassan Yousef, Sherif Sakr
EDBT4
2020 D-SmartML: A Distributed Automated Machine Learning Framework
abstract
Nowadays, machine learning is playing a crucial role in harnessing the value of massive data amount currently produced every day. The process of building a high-quality machine learning model is an iterative, complex and time-consuming process that requires solid knowledge about the various machine learning algorithms in addition to having a good experience with effectively tuning their hyper-parameters. With the booming demand for machine learning applications, it has been recognized that the number of knowledgeable data scientists can not scale with the growing data volumes and application needs in our digital world. Therefore, recently, several automated machine learning (AutoML) frameworks have been developed by automating the process of Combined Algorithm Selection and Hyper-parameter tuning (CASH). However, a main limitation of these frameworks is that they have been built on top of centralized machine learning libraries (e.g. scikit-learn) that can only work on a single node and thus they are not scalable to process and handle large data volumes. To tackle this challenge, we demonstrate D-SmartML, a distributed AutoML framework on top of Apache Spark, a distributed data processing framework. Our framework is equipped with a meta learning mechanism for automated algorithm selection and supports three different automated hyper-parameter tuning techniques: distributed grid search, distributed random search and distributed hyperband optimization. We will demonstrate the scalability of our framework on handling large datasets. In addition, we will show how our framework outperforms the-state-of-the-art framework for distributed AutoML optimization, TransmogrifAI.
Ahmed Abd Elrahman, Mohamed ElHelw, Radwa El Shawi, Sherif Sakr
ICDCS4
2020 Process Mining over Unordered Event Streams
abstract
Process mining is no longer limited to the one-off analysis of static event logs extracted from a single enterprise system. Rather, process mining may strive for immediate insights based on streams of events that are continuously generated by diverse information systems. This requires online algorithms that, instead of keeping the whole history of event data, work incrementally and update analysis results upon the arrival of new events. While such online algorithms have been proposed for several process mining tasks, from discovery through conformance checking to time prediction, they all assume that an event stream is ordered, meaning that the order of event generation coincides with their arrival at the analysis engine. Yet, once events are emitted by independent, distributed systems, this assumption may not hold true, which compromises analysis accuracy. In this paper, we provide the first contribution towards handling unordered event streams in process mining. Specifically, we formalize the notion of out-of-order arrival of events, where an online analysis algorithm needs to process events in an order different from their generation. Using directly-follows graphs as a basic model for many process mining tasks, we provide two approaches to handle such unorderedness, either through buffering or speculative processing. Our experiments with synthetic and real-life event data show that these techniques help mitigate the accuracy loss induced by unordered streams.
Ahmed Awad 0001, Matthias Weidlich 0001, Sherif Sakr
ICPM3
2020 On Teaching Web Stream Processing - Lessons Learned
Riccardo Tommasini 0001, Emanuele Della Valle, Marco Balduini, Sherif Sakr
ICWE4
2020 A First Step Towards a Streaming Linked Data Life-Cycle
Riccardo Tommasini 0001, Mohamed Ragab 0001, Alessandro Falcetta, Emanuele Della Valle, Sherif Sakr
ISWC (2)5
2020 Benchmarking big data systems: A survey
Fuad Bajaber, Sherif Sakr, Omar Batarfi, Abdulrahman H. Altalhi, Ahmed Barnawi
Comput. Commun.2
2020 The views, measurements and challenges of elasticity in the cloud: A review
Ahmed Barnawi, Sherif Sakr, Wenjing Xiao, Abdullah Al-Barakati
Comput. Commun.2
2019 ILIME: Local and Global Interpretable Model-Agnostic Explainer of Black-Box Decision
Radwa El Shawi, Youssef Sherif, Mouaz H. Al-Mallah, Sherif Sakr
ADBIS4
2019 D ^2 2 IA: Stream Analytics on User-Defined Event Intervals
Ahmed Awad 0001, Riccardo Tommasini 0001, Mahmoud Kamel, Emanuele Della Valle, Sherif Sakr
CAiSE5
2019 LDLCT An Instance-Based Framework for Lesion Detection on Lung CT Scans
abstract
Medical images have played a crucial role in transforming diagnostic medicine by providing the medical staff with several insights into the health status of every patient. Diagnosis of medical images is a very subjective process which is solely based on the physicians' expertise. In particular, lesion detection is a challenging task due to the various types, shapes and sizes of the lesions in the different organs. Thus, there is a crucial need to build computer-aided frameworks for automated analysis of medical images. In this paper, we present LDLCT, an instance-based framework for automated Lesion Detection on Lung CT Scans. The framework employs Restricted Botlzman Machines (RBM) network for feature extraction stage as an unsupervised feature mapping. The Random Forest (RF) classifier is used to distinguish between pixels from lesion and normal regions. Finally, a post-processing stage is implemented to filter out the false positive candidate lesions. In this study, we select 909 slices with 917 lesions from DeepLesion data set. LDLCT achieves lesion detection sensitivity of 89% with 5 false positives per image, the majority of them can be easily detected by the medical staff.
Tarun Khajuria, Eman Badr, Mouaz H. Al-Mallah, Sherif Sakr
CBMS4
2019 Interpretability in HealthCare A Comparative Study of Local Machine Learning Interpretability Techniques
abstract
Although complex machine learning models (e.g., Random Forest, Neural Networks) are commonly outperforming the traditional simple interpretable models (e.g., Linear Regression, Decision Tree), in the healthcare domain, clinicians find it hard to understand and trust these complex models due to the lack of intuition and explanation of their predictions. With the new General Data Protection Regulation (GDPR), the importance for plausibility and verifiability of the predictions made by machine learning models has become essential. To tackle this challenge, recently, several machine learning interpretability techniques have been developed and introduced. In general, the main aim of these interpretability techniques is to shed light and provide insights into the predictions process of the machine learning models and explain how the model predictions have resulted. However, in practice, assessing the quality of the explanations provided by the various interpretability techniques is still questionable. In this paper, we present a comprehensive experimental evaluation of three recent and popular local model agnostic interpretability techniques, namely, LIME, SHAP and Anchors on different types of real-world healthcare data. Our experimental evaluation covers different aspects for its comparison including identity, stability, separability, similarity, execution time and bias detection. The results of our experiments show that LIME achieves the lowest performance for the identity metric and the highest performance for the separability metric across all datasets included in this study. On average, SHAP has the smallest average time to output explanation across all datasets included in this study. For detecting the bias, SHAP enables the participants to better detect the bias.
Radwa El Shawi, Youssef Sherif, Mouaz H. Al-Mallah, Sherif Sakr
CBMS4
2019 Adaptive Watermarks: A Concept Drift-based Approach for Predicting Event-Time Progress in Data Streams
Ahmed Awad 0001, Jonas Traub, Sherif Sakr
EDBT3
2019 SmartML: A Meta Learning-Based Framework for Automated Selection and Hyperparameter Tuning for Machine Learning Algorithms
abstract
International audience
Mohamed Maher 0001, Sherif Sakr
EDBT2
2019 MINARET: A Recommendation Framework for Scientific Reviewers
abstract
International audience
Sherif Sakr, Mohamed Ragab 0001, Mohamed Maher 0001, Ahmed Awad 0001
EDBT1
2019 Editorial for Special issue of FGCS special issue on "Benchmarking big data systems"
Sherif Sakr, Albert Y. Zomaya, Athanasios V. Vasilakos
Future Gener. Comput. Syst.1
2018 HDM-MC in-Action: A Framework for Big Data Analytics across Multiple Clusters
abstract
Big data are increasingly collected and stored in a highly distributed infrastructures due to the development of several emerging technologies including sensor network, cloud computing, IoT and mobile computing among many other emerging technologies. In practice, the majority of existing big data processing frameworks (e.g., Hadoop, Spark, Flink) are designed based on the single-cluster setup with the assumptions of centralized management and homogeneous connectivity which makes them sub-optimal and sometimes infeasible to be applied for scenarios that require implementing data analytics jobs on highly distributed data sets (across racks, data centers or multi organizations). We demonstrate HDM-MC, a big data processing framework that is designed to enable the capability of performing large scale data analytics across multi-clusters with minimum extra overhead due to additional scheduling requirements. We describe the architecture and realization of the system using a step-by-step example scenario.
Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Sung Une Lee, Huijun Wu 0001
ICDCS2
2018 HDM: A Composable Framework for Big Data Processing
abstract
Over the past years, frameworks such as MapReduce and Spark have been introduced to ease the task of developing big data programs and applications. However, the jobs in these frameworks are roughly defined and packaged as executable jars without any functionality being exposed or described. This means that deployed jobs are not natively composable and reusable for subsequent development. Besides, it also hampers the ability for applying optimizations on the data flow of job sequences and pipelines. In this paper, we present the Hierarchically Distributed Data Matrix (HDM) which is a functional, strongly-typed data representation for writing composable big data applications. Along with HDM, a runtime framework is provided to support the execution, integration and management of HDM applications on distributed infrastructures. Based on the functional data dependency graph of HDM, multiple optimizations are applied to improve the performance of executing HDM jobs. The experimental results show that our optimizations can achieve improvements between 10 to 40 percent of the Job-Completion-Time for different types of applications when compared with the current state of art, Apache Spark.
Dongyao Wu, Liming Zhu 0001, Qinghua Lu 0001, Sherif Sakr
IEEE Trans. Big Data4
2018 A Differentiated Caching Mechanism to Enable Primary Storage Deduplication in Clouds
abstract
Existing primary deduplication techniques either use inline caching to exploit locality in primary workloads or use post-processing deduplication to avoid the negative impact on I/O performance. However, neither of them works well in the cloud servers running multiple services for the following two reasons: First, the temporal locality of duplicate data writes varies among primary storage workloads, which makes it challenging to efficiently allocate the inline cache space and achieve a good deduplication ratio. Second, the post-processing deduplication does not eliminate duplicate I/O operations that write to the same logical block address as it is performed after duplicate blocks have been written. A hybrid deduplication mechanism is promising to deal with these problems. Inline fingerprint caching is essential to achieving efficient hybrid deduplication. In this paper, we present a detailed analysis of the limitations of using existing caching algorithms in primary deduplication in the cloud. We reveal that existing caching algorithms either perform poorly or incur significant memory overhead in fingerprint cache management. To address this, we propose a novel fingerprint caching mechanism that estimates the temporal locality of duplicates in different data streams and prioritizes the cache allocation based on the estimation. We integrate the caching mechanism and build a hybrid deduplication system. Our experimental results show that the proposed mechanism provides significant improvement for both deduplication ratio and overhead reduction.
Huijun Wu 0001, Chen Wang 0008, Yinjin Fu, Sherif Sakr, Kai Lu 0001, Liming Zhu 0001
IEEE Trans. Parallel Distributed Syst.4
2017 Towards Big Data Analytics across Multiple Clusters
abstract
Big data are increasingly collected and stored in a highly distributed infrastructures due to the development of sensor network, cloud computing, IoT and mobile computing among many other emerging technologies. In practice, the majority of existing big-data-processing frameworks (e.g., Hadoop and Spark) are designed based on the single-cluster setup with the assumptions of centralized management and homogeneous connectivity which makes them sub-optimal and sometimes infeasible to apply for scenarios that require implementing data analytics jobs on highly distributed data sets (across racks, data centers or multi-organizations). In order to tackle this challenge, we present HDM-MC, a multi-cluster big data processing framework which is designed to enable the capability of performing large scale data analytics across multi-clusters with minimum extra overhead due to additional scheduling requirements. In this paper, we present the architecture and realization of the system. In addition, we evaluate the performance of our framework in comparison to other state-of-art single cluster big data processing frameworks.
Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Huijun Wu 0001
CCGrid2
2017 HDM: Optimized Big Data Processing with Data Provenance
Dongyao Wu, Sherif Sakr, Liming Zhu 0001
EDBT2
2017 On business process monitoring using cross-flow coordination
Zakaria Maamar, Noura Faci, Mohamed Sellami, Khouloud Boukadi, Fadwa Yahya, Ahmed Barnawi, Sherif Sakr
Serv. Oriented Comput. Appl.7
2016 Big Data 2.0 Processing Systems: Taxonomy and Open Challenges
Fuad Bajaber, Radwa El Shawi, Omar Batarfi, Abdulrahman H. Altalhi, Ahmed Barnawi, Sherif Sakr
J. Grid Comput.6
2016 Network-based social coordination of business processes
Zakaria Maamar, Noura Faci, Sherif Sakr, Mohamed Boukhebouze, Ahmed Barnawi
Inf. Syst.3
2015 Composable and efficient functional big data processing framework
abstract
Over the past years, frameworks such as MapReduce and Spark have been introduced to ease the task of developing big data programs and applications. However, the jobs in these frameworks are roughly defined and packaged as executable jars without any functionality being exposed or described. This means that deployed jobs are not natively composable and reusable for subsequent development. Besides, it also hampers the ability for applying optimizations on the data flow of job sequences and pipelines. In this paper, we present the Hierarchically Distributed Data Matrix (HDM) which is a functional, strongly-typed data representation for writing composable big data applications. Along with HDM, a runtime framework is provided to support the execution of HDM applications on distributed infrastructures. Based on the functional data dependency graph of HDM, multiple optimizations are applied to improve the performance of executing HDM jobs. The experimental results show that our optimizations can achieve improvements of between 10% to 60% of the Job-Completion-Time for different types of operation sequences when compared with the current state of art, Apache Spark.
Dongyao Wu, Sherif Sakr, Liming Zhu 0001, Qinghua Lu 0001
IEEE BigData2
2015 Liquid Benchmarking: A Platform for Democratizing the Performance Evaluation Process
abstract
Performances evaluation, reproducibility and benchmarking represent crucial aspects for assessing the practical impact of research results in the computer science eld. In spite of all the benets (e.g., increasing impact, increasing visibility, improving the research quality) that can be gained from performing extensive experimental evaluation or providing reproducible software artifacts and detailed description of experimental setup, the required eort for achiev
Sherif Sakr, Amin Shafaat, Fuad Bajaber, Ahmed Barnawi, Omar Batarfi, Abdulrahman H. Altalhi
EDBT1
2015 DREAM: Distributed RDF Engine with Adaptive Query Planner and Minimal Communication
abstract
The Resource Description Framework (RDF) and SPARQL query language are gaining wide popularity and acceptance. In this paper, we present DREAM, a distributed and adaptive RDF system. As opposed to existing RDF systems, DREAM avoids partitioning RDF datasets and partitions only SPARQL queries. By not partitioning datasets, DREAM offers a general paradigm for different types of pattern matching queries, and entirely averts intermediate data shuffling (only auxiliary data are shuffled). Besides, by partitioning queries, DREAM presents an adaptive scheme, which automatically runs queries on various numbers of machines depending on their complexities. Hence, in essence DREAM combines the advantages of the state-of-the-art centralized and distributed RDF systems, whereby data communication is avoided and cluster resources are aggregated. Likewise, it precludes their disadvantages, wherein system resources are limited and communication overhead is typically hindering. DREAM achieves all its goals via employing a novel graph-based, rule-oriented query planner and a new cost model. We implemented DREAM and conducted comprehensive experiments on a private cluster and on the Amazon EC2 platform. Results show that DREAM can significantly outperform three related popular RDF systems.
Mohammad Hammoud, Dania Abed Rabbou, Reza Nouri, Amin Beheshti, Sherif Sakr
Proc. VLDB Endow.5
2015 A Framework for Consumer-Centric SLA Management of Cloud-Hosted Databases
abstract
Service Level Agreements (SLA) represent the contract which captures the agreed upon guarantees between a service provider and its customers. The specifications of existing service level agreements (SLA) for cloud services are not designed to flexibly handle even relatively straightforward performance and technical requirements of consumer applications. In this article, we present a novel approach for SLA-based management of cloud-hosted databases from the consumer perspective. We present an end-to-end framework for consumer-centric SLA management of cloud-hosted databases. The framework facilitates adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. In this framework, the SLA of the consumer applications are declaratively defined in terms of goals which are subjected to a number of constraints that are specific to the application requirements. The framework continuously monitors the application-defined SLA and automatically triggers the execution of necessary corrective actions (scaling out/in the database tier) when required. The framework is database platform-agnostic, uses virtualization-based database replication mechanisms, and requires zero source code changes of the cloud-hosted software applications. The experimental results demonstrate the effectiveness of our SLA-based framework in providing the consumer applications with the required flexibility for achieving their SLA requirements.
Liang Zhao 0009, Sherif Sakr, Anna Liu
IEEE Trans. Serv. Comput.2
2014 CDPort: A Framework of Data Portability in Cloud Platforms
abstract
One of the main advantages of the cloud computing paradigm is that it simplifies the time-consuming processes of hardware provisioning, hardware procurement and software deployment. Currently, we are witnessing a proliferation in the number of cloud-hosted applications. However, one of the important challenges of the cloud computing paradigm that may hurt the growth of this technology is the interoperability and portability between cloud platforms. The developers and cloud users may lock-in to the first cloud they choose, or face difficulties when they have to move their data or software from one cloud platform to another. In this paper, we focus on the challenge of data portability between different cloud-based data storage services. In particular, we propose a common data model and a standardized API for the new generation of cloud-based NoSQL databases. The initial implementation of our framework covers three of the most popular NoSQL systems, namely, Google Datastore, Amazon SimpleDB and MongoDB. However, our framework is designed in a flexible way that it can be easily extended to support other NoSQL systems. Furthermore, our framework is equipped with tools that support the conversion, transformation and exchange of the data which is stored on the supported NoSQL databases of the framework. Finally, we describe the design and the proof-of-concept implementation of our framework using a case study.
Ebtesam Ahmad Alomari, Ahmed Barnawi, Sherif Sakr
iiWAS3
2014 Hybrid query execution engine for large attributed graphs
Sherif Sakr, Sameh Elnikety, Yuxiong He
Inf. Syst.1
2013 Improving Availability of Cloud-Based Applications through Deployment Choices
abstract
Deployment choices are critical in determining the availability of applications running in a cloud. But choosing good deployment for various software application components into virtual machines is a challenging task because of potential sharing of components among applications and potential interference from multi-tenancy. This paper presents an approach for improving the availability guarantee of software applications by optimizing the availability, performance and monetary cost trade-offs of different deployment choices. Our approach explicitly considers different classes of application requests during the decision process. The results of our experimental evaluation show that the approach can effectively improve the availability guarantees with little or negligible increase in the performance and monetary cost of the deployment choice.
Jim Zhanwen Li, Qinghua Lu 0001, Liming Zhu 0001, Leonard J. Bass, Xiwei Xu 0001, Sherif Sakr, Paul L. Bannerman, Anna Liu
IEEE CLOUD6
2013 Incorporating Uncertainty into In-Cloud Application Deployment Decisions for Availability
abstract
Cloud consumers have a variety of deployment related techniques, such as auto-scaling policies and recovery strategies, for dealing with the uncertainties in the cloud. Uncertainties can be characterized as stochastic (such as failures, disasters, and workload spikes) and subjective (such as choice among various deployment options). Cloud consumers must consider both stochastic and subjective uncertainties. Analytic support for consumers in selecting appropriate techniques and setting the required parameters in the face of different types of uncertainty is currently limited. In this paper, we propose a set of application availability analysis models that capture subjective uncertainties in addition to stochastic uncertainties. We built and validated the models by using industry best practices on deployment, and actual commercial products for disaster recovery and live migration. Our results show that the models permit more informed and quantitative availability analysis than industry best practices under a wide range of scenarios.
Qinghua Lu 0001, Xiwei Xu 0001, Liming Zhu 0001, Leonard J. Bass, Jim Zhanwen Li, Sherif Sakr, Paul L. Bannerman, Anna Liu
IEEE CLOUD6
2013 Consumer-centric SLA manager for cloud-hosted databases
abstract
We present an end-to-end framework for consumer-centric SLA management of virtualized database servers. The framework facilitates adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. In this framework, the SLA of the consumer applications are declaratively defined in terms of goals which are subjected to a number of constraints that are specific to the application requirements. The framework continuously monitors the application-defined SLA and automatically triggers the execution of necessary corrective actions (scaling out/in the database tier) when required. The framework is database platform-agnostic, uses virtualization-based database replication mechanisms and requires zero source code changes of the cloud-hosted application.
Liang Zhao 0009, Sherif Sakr, Anna Liu
CIKM2
2013 Is Your Cloud-Hosted Database Truly Elastic?
abstract
Elasticity has been recognized as one of the most appealing features for users of cloud services. It represents the ability to dynamically and rapidly scale up or down the allocated computing resources on demand. In practice, it is difficult to understand the elasticity requirements of a given application and workload, and to assess if the elasticity provided by a cloud service will meet these requirements. In this experience paper, we take the position that a deep understanding of the capabilities of cloud-hosted database services is a crucial requirement for cloud users in order to bring forward the vision of deploying data-intensive applications on cloud platforms. We argue that it is important that cloud users become able to paint a comprehensive picture of the relationship between the capabilities of the different type of cloud database services, the application characteristics and workloads, and the geographical distribution of the application clients and the underlying database replicas. We discuss the current elasticity capabilities of the different categories of cloud database services and identify some of the main challenges for deploying a truly elastic database tier on cloud environments. Finally, we propose a benchmarking mechanism that can evaluate the elasticity capabilities of cloud database services in different application scenarios and workloads.
Sherif Sakr, Anna Liu
SERVICES1
2013 Modeling performance of a parallel streaming engine: bridging theory and costs
abstract
While data are growing at a speed never seen before, parallel computing is becoming more and more essential to process this massive volume of data in a timely manner. Therefore, recently, concurrent computations have been receiving increasing attention due to the widespread adoption of multi-core processors and the emerging advancements of cloud computing technology. The ubiquity of mobile devices, location services, and sensor pervasiveness are examples of new scenarios that have created the crucial need for building scalable computing platforms and parallel architectures to process vast amounts of generated streaming data. In practice, efficiently operating these systems is hard due to the intrinsic complexity of these architectures and the lack of a formal and in-depth knowledge of the performance models and the consequent system costs. The Actor Model theory has been presented as a mathematical model of con- current computation that had enormous success in practice and inspired a number of contemporary work in this area. Recently, the Storm system has been presented as a realization of the principles of the Actor Model theory in the context of the large scale processing of streaming data. In this paper, we present, to the best of our knowledge, the first set of models that formalize the performance characteristics of a practical distributed, parallel and fault-tolerant stream processing system that follows the Actor Model theory. In particular, we model the characteristics of the data flow, the data processing and the system management costs at a fine granularity within the different steps of executing a distributed stream processing job. Finally, we present an experimental validation of the described performance models using the Storm system.
Ivan Bedini, Sherif Sakr, Bart Theeten, Alessandra Sala, Peter Cogan
ICPE2
2012 SLA-Based and Consumer-centric Dynamic Provisioning for Cloud Databases
abstract
One of the main advantages of the cloud computing paradigm is that it simplifies the time-consuming processes of hardware provisioning, hardware purchasing and software deployment. Currently, we are witnessing a proliferation in the number of cloud-hosted applications with a tremendous increase in the scale of the data generated as well as being consumed by such applications. Cloud-hosted database systems powering these applications form a critical component in the software stack of these applications. Service Level Agreements (SLA) represent the contract which captures the agreed upon guarantees between a service provider and its customers. The specifications of existing service level agreement (SLA) for cloud services are not designed for flexibly handling even relatively straightforward performance and technical requirements of consumer applications. The concerns of consumers for cloud services regarding the SLA management of their hosted applications within the cloud environments will gain increasing importance as cloud computing becomes more pervasive. This paper introduces the notion, challenges and the importance of SLA-based provisioning and cost management for cloud-hosted databases from the consumer perspective. We present an end-to-end framework that acts as a middleware which resides between the consumer applications and the cloud-hosted databases. The aim of the framework is to facilitate adaptive and dynamic provisioning of the database tier of the software applications based on application-defined policies for satisfying their own SLA performance requirements, avoiding the cost of any SLA violation and controlling the monetary cost of the allocated computing resources. The experimental results demonstrate that SLA-based provisioning is more adequate for providing consumer applications the required flexibility in achieving their goals.
Sherif Sakr, Anna Liu
IEEE CLOUD1
2012 Application-Managed Replication Controller for Cloud-Hosted Databases
abstract
Data replication is a well-known strategy to achieve the availability, scalability and performance improvement goals in the data management world. However, the cost of maintaining several database replicas always strongly consistent is very high. The CAP theorem shows that a shared-data system can choose at most two out of three properties: consistency, availability, and tolerance to partitions. In practice, most of the cloud-based data management systems tend to overcome the difficulties of distributed replication by relaxing the consistency guarantees of the system. In particular, they implement various forms of weaker consistency models such as eventual consistency. This solution is accepted by many new Web 2.0 applications (e.g. social networks) which could be more tolerant with a wider window of data staleness (replication delay).However, unfortunately, there are no generic application-independent and consumer-centric mechanisms by which software applications can specify and manage to what extent inconsistencies can be tolerated. We introduce an adaptive framework for database replication at the middleware layer of cloud environments. The framework provides flexible mechanisms to enable software applications of keeping several database replicas (that can be hosted in different data centers) with different levels of service level agreements (SLA) for their data freshness. The experimental evaluation demonstrates the effectiveness of our framework in providing the software applications with the required flexibility to achieve and optimize their requirements in terms of overall system throughput, data freshness and invested monetary cost.
Liang Zhao 0009, Sherif Sakr, Anna Liu
IEEE CLOUD2
2012 G-SPARQL: a hybrid engine for querying large attributed graphs
abstract
We propose a SPARQL-like language, G-SPARQL, for querying attributed graphs. The language expresses types of queries which of large interest for applications which model their data as large graphs such as: pattern matching, reachability and shortest path queries. Each query can combine both of structural predicates and value-based predicates (on the attributes of the graph nodes and edges). We describe an algebraic compilation mechanism for our proposed query language which is extended from the relational algebra and based on the basic construct of building SPARQL queries, the Triple Pattern. We describe a hybrid Memory/Disk representation of large attributed graphs where only the topology of the graph is maintained in memory while the data of the graph is stored in a relational database. The execution engine of our proposed query language splits parts of the query plan to be pushed inside the relational database while the execution of other parts of the query plan are processed using memory-based algorithms, as necessary. Experimental results on real datasets demonstrate the efficiency and the scalability of our approach and show that our approach outperforms native graph databases by several factors.
Sherif Sakr, Sameh Elnikety, Yuxiong He
CIKM1
2012 Trade-Off Analysis of Elasticity Approaches for Cloud-Based Business Applications
Basem Suleiman, Sherif Sakr, Srikumar Venugopal, Wasim Sadiq
WISE2
2011 A Query Language for Analyzing Business Processes Execution
Amin Beheshti, Boualem Benatallah, Hamid R. Motahari Nezhad, Sherif Sakr
BPM4
2011 Design by Selection: A Reuse-Based Approach for Business Process Modeling
Ahmed Awad 0001, Sherif Sakr, Matthias Kunze 0001, Mathias Weske
ER2
2011 CloudDB AutoAdmin: Towards a Truly Elastic Cloud-Based Data Store
abstract
In this paper, we present the design and the architecture of the CloudDB AutoAdmin system which aims to fill the existing gaps between the provided cloud database services and the requirements of the consumer applications. In particular, it focuses on facilitating the job of the cloud database consumers in implementing database applications as distributed, scalable, and elastic services with a minimum effort on the side of the application developer and a limited footprint in the application code.
Sherif Sakr, Liang Zhao 0009, Hiroshi Wada, Anna Liu
ICWS1
2010 An efficient features-based processing technique for supergraph queries
abstract
Graphs are widely used for modeling complicated data such as social networks, chemical compounds, protein interactions, XML documents and multimedia databases. To be able to effectively understand and utilize any collection of graphs, a graph database that efficiently supports elementary querying mechanisms is crucially required. Supergraph query is an important type of graph queries which has many practical applications. Given a graph database D, the answer set of a supergraph query q is computed by retrieving all graphs in D which are fully contained in q. A primary challenge in computing the answers of graph queries is that pair-wise comparisons of graphs are usually hard problems. For example, subgraph isomorphism is known to be NP-complete. Clearly, the success of any graph database application is directly dependent on the efficiency of the graph indexing and query processing mechanisms. In this paper, we study the problem of using the relational infrastructure to achieve an efficient evaluation of supergraph queries. We rely on an effective and efficient layer of features-based summary structures, called graph features knowledge, to reduce the required number of pair-wise graph comparisons and boost the efficiency of query processing. Finally, we conduct an extensive set of experiments on real and synthetic data sets to demonstrate the efficiency and the scalability of our approach.
Sherif Sakr, Ghazi Al-Naymat
IDEAS1
2010 Efficient and Adaptable Query Workload-Aware Management for RDF Data
Hooran MahmoudiNasab, Sherif Sakr
WISE2
2010 A framework for querying graph-based business process models
abstract
We present a framework for querying and reusing graph-based business process models. The framework is based on a new visual query language for business processes called BPMN-Q. The language addresses processes definitions and extends the standard BPMN visual notations for modeling business processes for its concrete syntax. BPMN-Q is used to query process models by matching a process model graph to a query graph. Moreover, the reusing framework is enhanced with a semantic query expander component. This component provides the users with the flexibility to get not only the perfectly matched process models to their queries but also the models with high similarity. The query engine of the framework is built on top of traditional RDBMS. A novel decomposition based and selectivity-aware relational processing mechanism is employed to achieve an efficient and scalable performance for graph-based BPMN-Q queries.
Sherif Sakr, Ahmed Awad 0001
WWW1
2010 Efficient Relational Techniques for Processing Graph Queries
Sherif Sakr, Ghazi Al-Naymat
J. Comput. Sci. Technol.1
2009 GraphREL: A Decomposition-Based and Selectivity-Aware Relational Framework for Processing Sub-graph Queries
Sherif Sakr
DASFAA1
2009 FeedRank: A Semantic-Based Management System of Web Feeds
Hooran MahmoudiNasab, Sherif Sakr
IDEAL2
2009 XML compression techniques: A survey and comparison
Sherif Sakr
J. Comput. Syst. Sci.1
2009 Cardinality-Aware Purely Relational XQuery Processor
abstract
Recently, the use of eXtensible Markup Language (XML) continues to grow in popularity, large repositories of XML documents are going to emerge, and users are likely to pose increasingly more complex queries on these data sets. In 2001 XQuery is decided by the World Wide Web Consortium (W3C) as the standard XML query language. In this article, we describe the design and implementation of an efficient and scalable purely relational XQuery processor which translates expressions of the XQuery language into their equivalent SQL evaluation scripts. The experiments of this article demonstrated the efficiency and scalability of our purely relational approach in comparison to the native XML/XQuery functionality supported by conventional RDBMSs and has shown that our purely relational approach for implementing XQuery processor deserves to be pursued further.
Sherif Sakr
J. Database Manag.1
2008 XSelMark: A Micro-benchmark for Selectivity Estimation Approaches of XML Queries
Sherif Sakr
DEXA1
2008 Improving the Relational Evaluation of XML Queries by Means of Path Summaries
Sherif Sakr
IDEAL1
2008 Dependable cardinality forecasts for XQuery
abstract
Though inevitable for effective cost-based query rewriting, the derivation of meaningful cardinality estimates has remained a notoriously hard problem in the context of XQuery. By basing the estimation on a relational representation of the XQuery syntax, we show how existing cardinality estimation techniques for XPath and proven relational estimation machinery can play together to yield dependable forecasts for arbitrary XQuery (sub)expressions. Our approach benefits from a light-weight form of data flow analysis. Abstract domain identifiers guide our query analyzer through the estimation process and allow for informed decisions even in case of deeply nested XQuery expressions. A variant of projection paths [15] provides a versatile interface into which existing techniques for XPath cardinality estimation can be plugged in seamlessly. We demonstrate an implementation of this interface based on data guides. Experiments show how our approach can equally cope with both, structure-and value-based queries. It is robust with respect to intermediate estimation errors, from which we typically found our implementation to recover gracefully.
Jens Teubner, Torsten Grust, Sebastian Maneth, Sherif Sakr
Proc. VLDB Endow.4
2007 A SQL: 1999 code generator for the pathfinder xquery compiler
abstract
The Pathfinder XQuery compiler has been enhanced by a new code generator that can target any SQL:1999-compliant relational database system(RDBMS). This code generator marks an important next step towards truly relational XQuery processing, a branch of database technology that aims to turn RDBMSs into highly efficient XML and XQuery processors without the need to invade the relational database kernel. Pathfinder, a retargetable front-end compiler, translates input XQuery expressions into DAG-shaped relational algebra plans. The code generator then turns these plans into sequences of either SQL:1999 statements or view definitions which jointly implement the (sometimes intricate) XQuery semantics. In a sense, this demonstration thus lets relational algebra and SQL swap their traditional roles in database query processing. The result is a code generator that (1) supports an almost complete dialect of XQuery, (2) can target any RDBMS with a SQL:1999 language interface, and (3) exhibits quite promising performance characteristics when run against high-volume XML data as well as complex XQuery expressions.
Torsten Grust, Manuel Mayr, Jan Rittinger, Sherif Sakr, Jens Teubner
SIGMOD Conference4
2004 XQuery on SQL Hosts
Torsten Grust, Sherif Sakr, Jens Teubner
VLDB2