EDBT 2026 Demo / reviewers in the wild / expert
Mohamed A. Sharaf
dblp:s/MohamedASharaf
· DBLP profile ↗
54ranked-venue papers
12as first author
4since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 44 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 1 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
13 papers |
Information retrieval · 51% Query processing and optimization · 16% Data stream processing · 9% | |
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Storage systems · 46% Cloud and datacenter computing · 32% Distributed systems · 10% | |
| Computer graphics and multimedia
1 paper |
Visualization and visual analytics · 50% Computational photography and imaging · 50% |
Topics — the 30 heaviest of 39, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
query formulation |
0.6 | 2 | 2020 | Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020 CoRE: A Context-Aware RelationExtraction Method for Relation Completion · IEEE Trans. Knowl. Data Eng. 2014 |
Information retrieval › query formulation › query generation
boolean query generation |
0.4 | 1 | 2020 | Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020 |
Information retrieval › ranking
learning to rank |
0.4 | 1 | 2020 | Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020 |
Information retrieval
retrieval models |
0.4 | 1 | 2020 | Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020 |
Information retrieval
search engines |
0.4 | 1 | 2020 | Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020 |
Data integration and cleaning › missing data
missing value imputation |
0.4 | 1 | 2019 | WebPut: A Web-Aided Data Imputation System for the General Type of Missing String Attribute Values · ICDE 2019 |
Computational photography and imaging
view recommendation |
0.3 | 1 | 2018 | Efficient Recommendation of Aggregate Data Visualizations · IEEE Trans. Knowl. Data Eng. 2018 |
Visualization and visual analytics
visualization recommendation |
0.3 | 1 | 2018 | Efficient Recommendation of Aggregate Data Visualizations · IEEE Trans. Knowl. Data Eng. 2018 |
Recommender systems › domain-specific recommendation
visualization recommendation |
0.2 | 1 | 2016 | MuVE: Efficient Multi-Objective View Recommendation for Visual Data Exploration · ICDE 2016 |
Data stream processing › continuous query processing
continuous query scheduling |
0.2 | 3 | 2008 | Algorithms and metrics for processing multiple heterogeneous continuous queries · ACM Trans. Database Syst. 2008 Scheduling continuous queries in data stream management systems · Proc. VLDB Endow. 2008 Efficient Scheduling of Heterogeneous Continuous Queries · VLDB 2006 |
Query processing and optimization
multi-query optimization |
0.2 | 2 | 2012 | Three-Level Processing of Multiple Aggregate Continuous Queries · ICDE 2012 Algorithms and metrics for processing multiple heterogeneous continuous queries · ACM Trans. Database Syst. 2008 |
Information retrieval
search result diversification |
0.2 | 1 | 2015 | Progressive diversification for column-based data exploration platforms · ICDE 2015 |
Natural language and speech › Information extraction and text analysis
relation extraction |
0.2 | 1 | 2014 | CoRE: A Context-Aware RelationExtraction Method for Relation Completion · IEEE Trans. Knowl. Data Eng. 2014 |
Data mining › dimensionality reduction
feature selection |
0.2 | 1 | 2014 | Mining Personal Health Index from Annual Geriatric Medical Examinations · ICDM 2014 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.2 | 1 | 2014 | AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014 |
Storage systems
key-value storage |
0.2 | 1 | 2014 | AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014 |
Storage systems › key-value storage
NoSQL database |
0.2 | 1 | 2014 | AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014 |
Data stream processing › continuous query processing
continuous aggregate query |
0.1 | 1 | 2012 | Three-Level Processing of Multiple Aggregate Continuous Queries · ICDE 2012 |
Transaction processing and concurrency control
transaction scheduling |
0.1 | 1 | 2009 | Optimizing i/o-intensive transactions in highly interactive applications · SIGMOD Conference 2009 |
Storage systems
i/o scheduling |
0.1 | 1 | 2009 | Optimizing i/o-intensive transactions in highly interactive applications · SIGMOD Conference 2009 |
Indexing and storage engines
buffer management |
0.1 | 1 | 2008 | Dynamic partitioning of the cache hierarchy in shared data centers · Proc. VLDB Endow. 2008 |
Query processing and optimization › multi-query optimization
query plan sharing |
0.1 | 1 | 2008 | Algorithms and metrics for processing multiple heterogeneous continuous queries · ACM Trans. Database Syst. 2008 |
Data stream processing
stream processing systems |
0.1 | 1 | 2008 | Scheduling continuous queries in data stream management systems · Proc. VLDB Endow. 2008 |
Memory systems › cache management
cache partitioning |
0.1 | 1 | 2008 | Dynamic partitioning of the cache hierarchy in shared data centers · Proc. VLDB Endow. 2008 |
Cloud and datacenter computing
resource management |
0.1 | 1 | 2008 | Dynamic partitioning of the cache hierarchy in shared data centers · Proc. VLDB Endow. 2008 |
Data mining › exploratory data analysis
visual exploration |
0.1 | 1 | 2016 | MuVE: Efficient Multi-Objective View Recommendation for Visual Data Exploration · ICDE 2016 |
Distributed systems
replication |
0.1 | 1 | 2014 | AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014 |
Internet of things and sensor networks › wireless sensor network
in-network aggregation |
0.0 | 1 | 2004 | Balancing energy efficiency and quality of aggregate data in sensor networks · VLDB J. 2004 |
Cellular and mobile networks
scalable information dissemination |
0.0 | 1 | 2003 | An Optimized Multicast-based Data Dissemination Middleware · ICDE 2003 |
Distributed systems › distributed communication
data dissemination |
0.0 | 1 | 2003 | An Optimized Multicast-based Data Dissemination Middleware · ICDE 2003 |
Methods — techniques the papers use, named apart from their topics
query variation sampling · 0.4learning to rank · 0.4crowdsourcing · 0.4pattern-based extraction · 0.4model selection · 0.4data preprocessing · 0.4context term learning · 0.4pruning · 0.3multi-objective optimization · 0.3multi-objective utility function · 0.2incremental pruning · 0.2partial distance computation · 0.2i/o throughput optimization · 0.2deadline-aware scheduling · 0.2quality-of-service scheduling · 0.2quality-of-data management · 0.2tardiness minimization · 0.1adaptive scheduling · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | From Incomplete Data to Accurate Aggregation: Exploring the Accuracy-Efficiency Trade-off of Data Imputation MethodsabstractMissing data remains a critical challenge in data analytics, especially for aggregation-based descriptive tasks. In addition to evaluating the accuracy of data imputation methods at the cell-level, several works have studied their impact on downstream predictive analytics. However, the interplay between data imputation and data aggregation remains largely underexplored. To address that limitation, in this work, we evaluate the impact of data imputation on the accuracy of aggregate queries. Our evaluation is conducted on several real datasets, measuring the accuracy at both the cell-level and the aggregatelevel, as well as runtime cost. Our results show clear trade-offs between accuracy and efficiency, providing guidance for selecting imputation strategies in aggregation-based analytical pipelines. Linda Mohammed, Heba Helal, Mohamed A. Sharaf |
AICCSA | 3 |
| 2023 | Strategies for Optimizing Time Series Visual Data ExplorationabstractDealing with high-dimensional time series data makes the process of recommending visualizations with "interesting" insights difficult. The challenge originates from finding a way to obtain the recommended visualizations efficiently without compromising their quality. Identifying such visualizations manually is considered a labor-intensive and time-consuming process. In response, this paper introduces different techniques designed to optimize the automated recommendation process. These techniques are entirely based on the concept of computation sharing and pruning. Furthermore, we provide a glimpse into our future research works in PhD thesis. The objective is to broaden the scope of our current work and enhance the generality of our problem statement. Heba Helal, Mohamed A. Sharaf |
AICCSA | 2 |
| 2023 | TiVEx: Optimized Processing for Time Series Visual ExplorationabstractTo facilitate fast-visual data analysis, there is a need for recommending top-k views with "interesting" insights automatically. However, working with high-dimensional time series data makes the process of view recommendations difficult. The primary obstacle lies in finding an automatic way to generate views with less processing time (efficiency) while still closely aligning with the ground truth (effectiveness). In this paper, we propose TiVEx (Time Series Visual Exploration), a technique to address this challenge. TiVEx aims to achieve a balance between efficiency and effectiveness in generating view recommendations. Through extensive experiments, we demonstrate significant cost savings achieved by TiVEx, indicating its efficiency. Furthermore, our analysis delves into the exploration of striking the right balance between efficiency and effectiveness. Heba Helal, Mohamed A. Sharaf, Mohammad M. Masud 0001, Panos K. Chrysanthis |
AICCSA | 2 |
| 2022 | CovidLens: Visually Understanding the Covid-19 Indicators through the Lens of Mobility DataabstractSince the onset of the Covid-19 pandemic, an over-whelming amount of related data has been released. In an attempt to gain insights from that data, multiple public data visualization dashboards have been deployed. Differently from such dashboards, which mainly support basic data filtering and visualization of separate datasets, in this work, we propose CovidLens, which: 1) integrates various Covid-19 indicators and is centred around the Google Community Mobility Report dataset, 2) supports similarity search for finding similar and correlated patterns and trends across the integrated datasets, and 3) automatically recommends insightful visualizations that unlocks valuable insights into the pandemic effects. To that end, we will be presenting the employed dataset, together with the design, implementation, and multiple usage scenarios of our proposed CovidLens. Mohamed A. Sharaf, Xiaozhong Zhang, Panos K. Chrysanthis, Wadima Alsaedi, Maitha Alkalbani, Heba Helal, Alyazia Aldhaheri |
MDM | 1 |
| 2020 | Quality Matters: Understanding the Impact of Incomplete Data on Visualization Recommendation
Rischan Mafrur, Mohamed A. Sharaf, Guido Zuccon |
DEXA (1) | 2 |
| 2020 | Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic ReviewsabstractSearching medical literature for synthesis in a systematic review is a complex and labour intensive task. In this context, expert searchers construct lengthy Boolean queries. The universe of possible query variations can be massive: a single query can be composed of hundreds of field-restricted search terms/phrases or ontological concepts, each grouped by a logical operator nested to depths of sometimes five or more levels deep. With the many choices about how to construct a query, it is difficult to both formulate and recognise effective queries. To address this challenge, automatic methods have recently been explored for generating and selecting effective Boolean query variations for systematic reviews. The limiting factor of these methods is that it is computationally infeasible to process all query variations for training the methods. To overcome this, we propose novel query variation sampling methods for training Learning to Rank models to rank queries. Our results show that query sampling methods do directly impact the ability of a Learning to Rank model to effectively identify good query variations. Thus, selecting appropriate query sampling methods is a key problem for the automatic reformulation of effective Boolean queries for systematic review literature search. We find that the best sampling strategies are those which balance the diversity of queries with the quantity of queries. Harrisen Scells, Guido Zuccon, Mohamed A. Sharaf, Bevan Koopman |
WWW | 3 |
| 2020 | Serendipity-based Points-of-Interest NavigationabstractTraditional venue and tour recommendation systems do not necessarily provide a diverse set of recommendations and leave little room for serendipity . In this article, we design MPG, a Mobile Personal Guide that recommends: (i) a set of diverse yet surprisingly interesting venues that are aligned to user preferences and (ii) a set of routes, constructed from the recommended venues. We also introduce EPUI, an Experimental Platform for Urban Informatics. Our comparison with the state-of-the-art schemes indicates that MPG is capable of providing high-quality venues and route recommendations while incorporating seamlessly both the notion of diversity and that of serendipity. Xiaoyu Ge, Panos K. Chrysanthis, Konstantinos Pelechrinis, Demetris Zeinalipour, Mohamed A. Sharaf |
ACM Trans. Internet Techn. | 5 |
| 2019 | WebPut: A Web-Aided Data Imputation System for the General Type of Missing String Attribute ValuesabstractIn this demonstration, we present an end-to-end web-aided data imputation prototype system named WebPut. WebPut consults the Web for imputing the missing values in a local database when the traditional inferring-based imputation method has difficulties in getting the right answers. Specifically, WebPut investigates the interaction between the local inferring-based imputation methods and the web-based retrieving methods and shows that retrieving a small number of selected missing values can greatly improve the imputation recall of the inferring-based methods. Besides, WebPut also incorporates a crowd intervention component that can get advice from humans in case that the web-based imputation methods may have difficulties in making the right decisions. We demonstrate, step by step, how WebPut fills an incomplete table with each of its components. Shuangli Shan, Zhixu Li, Qiang Yang 0015, Jia Zhu 0003, Mohamed A. Sharaf, Xiaofang Zhou 0001 |
ICDE | 6 |
| 2018 | DiVE: Diversifying View Recommendation for Visual Data ExplorationabstractTo support effective data exploration, there has been a growing interest in developing solutions that can automatically recommend data visualizations that reveal interesting and useful data-driven insights. In such solutions, a large number of possible data visualization views are generated and ranked according to some metric of importance (e.g., a deviation-based metric), then the top-k most important views are recommended. However, one drawback of that approach is that it often recommends similar views, leaving the data analyst with a limited amount of gained insights. To address that limitation, in this work we posit that employing diversification techniques in the process of view recommendation allows eliminating that redundancy and provides a good and concise coverage of the possible insights to be discovered. To that end, we propose a hybrid objective utility function, which captures both the importance, as well as the diversity of the insights revealed by the recommended views. While in principle, traditional diversification methods (e.g., Greedy Construction) provide plausible solutions under our proposed utility function, they suffer from a significantly high query processing cost. In particular, directly applying such methods leads to a "process-first-diversify-next" approach, in which all possible data visualization are generated first via executing a large number of aggregate queries. To address that challenge, we propose an integrated scheme called DiVE, which efficiently selects the top-k recommended view based on our hybrid utility function. DiVE leverages the properties of both the importance and diversity metrics to prune a large number of query executions without compromising the quality of recommendations. Our experimental evaluation on real datasets shows the performance gains provided by DiVE. Rischan Mafrur, Mohamed A. Sharaf, Hina A. Khan |
CIKM | 2 |
| 2018 | Efficient Recommendation of Aggregate Data VisualizationsabstractData visualization is a common and effective technique for data exploration. However, for complex data, it is infeasible for an analyst to manually generate and browse all possible visualizations for insights. This observation motivated the need for automated solutions that can effectively recommend such visualizations. The main idea underlying those solutions is to evaluate the utility of all possible visualizations and then recommend the top-k visualizations. This process incurs high data processing cost, that is further aggravated by the presence of numerical dimensional attributes. To address that challenge, we propose novel view recommendation schemes, which incorporate a hybrid multi-objective utility function that captures the impact of numerical dimension attributes. Our first scheme, Multi-Objective View Recommendation for Data Exploration (MuVE), adopts an incremental evaluation of our multi-objective utility function, which allows pruning of a large number of low-utility views and avoids unnecessary objective evaluations. Our second scheme, upper MuVE (uMuVE), further improves the pruning power by setting the upper bounds on the utility of views and allowing interleaved processing of views, at the expense of increased memory usage. Finally, our third scheme, Memory-aware uMuVE (MuMuVE), provides pruning power close to that of uMuVE, while keeping memory usage within a pre-specified limit. Humaira Ehsan, Mohamed A. Sharaf, Panos K. Chrysanthis |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2017 | In Search for Relevant, Diverse and Crowd-screen Points of InterestsabstractIn this demo we present a prototype of an experimental platform for evaluating item recommendation algorithms. The application domain for our system is that of digital city guides. Our prototype implementation allows the user to explore different algorithms and compare their output. Among the algorithms implemented is MPG, which aims at providing a diverse set of recommendations better aligned with user preferences. MPG takes into consideration the user preferences (e.g., reach willing to cover, types of venues interested in exploring etc.), the popularity of the establishments as well as their distance from the current location of the user by combining them into a single composite score. We provide a web interface, which outputs on a map the recommended locations along with metadata (e.g., type and name of location, relevance and diversity scores, etc.). It also illustrates the potential of the Preferential Diversity approach on which MPG is based. Xiaoyu Ge, Samanvoy Panati, Konstantinos Pelechrinis, Panos K. Chrysanthis, Mohamed A. Sharaf |
EDBT | 5 |
| 2017 | Model-Based Diversification for Sequential Exploratory QueriesabstractToday, data exploration platforms are widely used to assist users in locating interesting objects within large volumes of scientific and business data. In those platforms, users try to make sense of the underlying data space by iteratively posing numerous queries over large databases. While diversification of query results, like other data summarization techniques, provides users with quick insights into the huge query answer space, it adds additional complexity to an already computationally expensive data exploration task. To address this challenge, in this paper we propose a diversification scheme that targets the problem of efficiently diversifying the results of multiple queries within and across different data exploratory sessions. Our proposed scheme relies on a model-based diversification method and an ordered cache. In particular, we employ an adaptive regression model to estimate the diversity of a diverse subset. Such estimation of diversity value allows us to select diverse results without scanning all the query results. In order to further expedite the diversification process, we propose an order-based caching scheme to leverage the overlap between sequence of data exploration queries. Our extensive experimental evaluation on both synthetic and real data sets shows the significant benefits provided by our scheme as compared to the existing methods. Hina A. Khan, Mohamed A. Sharaf |
Data Sci. Eng. | 2 |
| 2017 | Efficient schemes for similarity-aware refinement of aggregation queries
Abdullah M. Albarrak, Mohamed A. Sharaf |
World Wide Web | 2 |
| 2016 | REQUEST: A scalable framework for interactive construction of exploratory queriesabstractExploration over large datasets is a key first step in data analysis, as users may be unfamiliar with the underlying database schema and unable to construct precise queries that represent their interests. Such data exploration task usually involves executing numerous ad-hoc queries, which requires a considerable amount of time and human effort. In this paper, we present REQUEST, a novel framework that is designed to minimize the human effort and enable both effective and efficient data exploration. REQUEST supports the query-from-examples style of data exploration by integrating two key components: 1) Data Reduction, and 2) Query Selection. As instances of the REQUEST framework, we propose several highly scalable schemes, which employ active learning techniques and provide different levels of efficiency and effectiveness as guided by the user's preferences. Our results, on real-world datasets from Sloan Digital Sky Survey, show that our schemes on average require 1-2 orders of magnitude fewer feedback questions than the random baseline, and 3-16× fewer questions than the state-of-the-art, while maintaining interactive response time. Moreover, our schemes are able to construct, with high accuracy, queries that are often undetectable by current techniques. Xiaoyu Ge, Yanbing Xue, Mohamed A. Sharaf, Panos K. Chrysanthis |
IEEE BigData | 4 |
| 2016 | MuVE: Efficient Multi-Objective View Recommendation for Visual Data ExplorationabstractTo support effective data exploration, there is a well-recognized need for solutions that can automatically recommend interesting visualizations, which reveal useful insights into the analyzed data. However, such visualizations come at the expense of high data processing costs, where a large number of views are generated to evaluate their usefulness. Those costs are further escalated in the presence of numerical dimensional attributes, due to the potentially large number of possible binning aggregations, which lead to a drastic increase in the number of possible visualizations. To address that challenge, in this paper we propose the MuVE scheme for Multi-Objective View Recommendation for Visual Data Exploration. MuVE introduces a hybrid multi-objective utility function, which captures the impact of binning on the utility of visualizations. Consequently, novel algorithms are proposed for the efficient recommendation of data visualizations that are based on numerical dimensions. The main idea underlying MuVE is to incrementally and progressively assess the different benefits provided by a visualization, which allows an early pruning of a large number of unnecessary operations. Our extensive experimental results show the significant gains provided by our proposed scheme. Humaira Ehsan, Mohamed A. Sharaf, Panos K. Chrysanthis |
ICDE | 2 |
| 2015 | Progressive diversification for column-based data exploration platformsabstractIn Data Exploration platforms, diversification has become an essential method for extracting representative data, which provide users with a concise and meaningful view of the results to their queries. However, the benefits of diversification are achieved at the expense of an additional cost for the post-processing of query results. For high dimensional large result sets, the cost of diversification is further escalated due to massive distance computations required to evaluate the similarity between results. To address that challenge, in this paper we propose the Progressive Data Diversification (pDiverse) scheme. The main idea underlying pDiverse is to utilize partial distance computation to reduce the amount of processed data. Our extensive experimental results on both synthetic and real data sets show that our proposed scheme outperforms existing diversification methods in terms of both I/O and CPU costs. Hina A. Khan, Mohamed A. Sharaf |
ICDE | 2 |
| 2015 | Emerging event detection in social networks with location sensitivity
Sayan Unankard, Xue Li 0001, Mohamed A. Sharaf |
World Wide Web | 3 |
| 2014 | SAQR: An Efficient Scheme for Similarity-Aware Query Refinement
Abdullah M. Albarrak, Mohamed A. Sharaf, Xiaofang Zhou 0001 |
DASFAA (1) | 2 |
| 2014 | AQUAS: A quality-aware scheduler for NoSQL data storesabstractNoSQL key-value data stores provide an attractive solution for big data management. With the help of data partitioning and replication, those data stores achieve higher levels of availability, scalability and reliability. Such design choices typically exhibit a tradeoff in which data freshness is sacrificed in favor of reduced access latency. At the replica-level, this tradeoff is primarily shaped by the resource allocation strategies deployed for managing the processing of user queries and replica updates. In this demonstration, we showcase AQUAS: a quality-aware scheduler for Cassandra, which allows application developers to specify requirements on quality of service (QoS) and quality of data (QoD). AQUAS efficiently allocates the available replica resources to execute the incoming read/write tasks so that to minimize the penalties incurred by violating those requirements. We demonstrate AQUAS based on our implementation of a microblogging system. Chen Xu 0001, Mohamed A. Sharaf, Minqi Zhou, Aoying Zhou |
ICDE | 3 |
| 2014 | Mining Personal Health Index from Annual Geriatric Medical ExaminationsabstractPeople take regular medical examinations mostly not for discovering diseases but for having a peace of mind regarding their health status. Therefore, it is important to give them an overall feedback with respect to all the health indicators that have been ranked against the whole population. In this paper, we propose a framework of mining Personal Health Index (PHI) from a large and comprehensive geriatric medical examination (GME) dataset. We define PHI as an overall score of personal health status based on a complement probability of health risks. The health risks are calculated using the information from the cause of death (COD) dataset that is linked to the GME dataset. Especially, the highest health risk is revealed in the cases of people who had been taking GME for some years and then passed away for medical reasons. The proposed framework consists of methods in data pre-processing, feature extraction and selection, and model selection. The effectiveness of the proposed framework is validated by a set of comprehensive experiments based on the records of 102,258 participants. As the first of this kind, our work provides a baseline for further research. Ling Chen 0004, Xue Li 0001, Sen Wang 0001, Hsiao-Yun Hu, Nicole Huang, Quan Z. Sheng, Mohamed A. Sharaf |
ICDM | 7 |
| 2014 | ORange: Objective-Aware Range Query RefinementabstractIn this demo paper we present Orange, a system prototype for objective-aware range query refinement. Orange essentially refines a range query to meet a pre-specified cardinality constraint while taking into account the (dis)similarity between the initial query and its corresponding refined version. To achieve this goal, Orange employes the novel scheme SAQR for efficient similarity-aware query refinement. The main idea underlying SAQR is to utilize the pre-defined constraints on cardinality and similarity in order to bound the search space and quickly find a refined query, which meets the user's expectations. We showcase Orange in a web-based application which aims to guide planners in allocating service zones for police patrol units using real and historical dataset of crime incidents. Abdullah M. Albarrak, Tatiana Noboa, Hina A. Khan, Mohamed A. Sharaf, Xiaofang Zhou 0001, Shazia Sadiq |
MDM (1) | 4 |
| 2014 | Efficient Retrieval of Top-K Most Similar Users from Travel Smart Card DataabstractUnderstanding the dynamics of human daily mobility patterns is essential for the management and planning of urban facilities and services. Travel smart cards, which record users' public transporting histories, capture rich information of users' mobility pattern. This provides the opportunity to discover valuable knowledge from these transaction records. In recent years, research on measuring user similarity for behavior analysis has attracted a lot of attention in applications such as recommendation systems, crowd behavior analysis applications, and numerous data mining tasks. In this paper, our goal is to estimate the similarity between users' travel patterns according to their travel smart card data. The core of our proposal is a novel user similarity measurement, namely, Travel Spatial-Temporal Similarity (TST), which measures the spatial range and temporal similarity between users. Moreover, we also propose a hybrid index structure, which integrates inverted files and cluster-based partitioning, to allow for efficient retrieval of the top-K most similar users. Through experimental evaluation, our proposed approach is shown to deliver scalable performance. Bolong Zheng, Kai Zheng 0001, Mohamed A. Sharaf, Xiaofang Zhou 0001, Shazia Sadiq |
MDM (1) | 3 |
| 2014 | DivIDE: efficient diversification for interactive data explorationabstractToday, Interactive Data Exploration (IDE) has become a main constituent of many discovery-oriented applications, in which users repeatedly submit exploratory queries to identify interesting subspaces in large data sets. Returning relevant yet diverse results to such queries provides users with quick insights into a rather large data space. Meanwhile, search results diversification adds additional cost to an already computationally expensive exploration process. To address this challenge, in this paper, we propose a novel diversification scheme called DivIDE, which targets the problem of efficiently diversifying the results of queries posed during data exploration sessions. In particular, our scheme exploits the properties of data diversification functions while leveraging the natural overlap occurring between the results of different queries so that to provide significant reductions in processing costs. Our extensive experimental evaluation on both synthetic and real data sets shows the significant benefits provided by our scheme as compared to existing methods. Hina A. Khan, Mohamed A. Sharaf, Abdullah M. Albarrak |
SSDBM | 2 |
| 2014 | Predicting Elections from Social Networks Based on Sub-event Detection and Sentiment Analysis
Sayan Unankard, Xue Li 0001, Mohamed A. Sharaf |
WISE (2) | 3 |
| 2014 | Quality-aware schedulers for weak consistency key-value data stores
Chen Xu 0001, Mohamed A. Sharaf, Xiaofang Zhou 0001, Aoying Zhou |
Distributed Parallel Databases | 2 |
| 2014 | A framework for data quality aware query systems
Naiem Khodabandehloo Yeganeh, Shazia Sadiq, Mohamed A. Sharaf |
Inf. Syst. | 3 |
| 2014 | CoRE: A Context-Aware RelationExtraction Method for Relation CompletionabstractWe identify Relation Completion (RC) as one recurring problem that is central to the success of novel big data applications such as Entity Reconstruction and Data Enrichment.Given a semantic relation R, RC attempts at linking entity pairs between two entity lists under the relation R. To accomplish the RC goals, we propose to formulate search queries for each query entity α based on some auxiliary information, so that to detect its target entity β from the set of retrieved documents.For instance, a Pattern-based method (PaRE) uses extracted patterns as the auxiliary information in formulating search queries.However, high-quality patterns may decrease the probability of finding suitable target entities.As an alternative, we propose CoRE method that uses context terms learned surrounding the expression of a relation as the auxiliary information in formulating queries.The experimental results based on several real-world web data collections demonstrate that CoRE reaches a much higher accuracy than PaRE for the purpose of RC. Zhixu Li, Mohamed A. Sharaf, Laurianne Sitbon, Xiaoyong Du 0001, Xiaofang Zhou 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2014 | A web-based approach to data imputation
Zhixu Li, Mohamed A. Sharaf, Laurianne Sitbon, Shazia Sadiq, Marta Indulska, Xiaofang Zhou 0001 |
World Wide Web | 2 |
| 2013 | Location-Based Emerging Event Detection in Social Networks
Sayan Unankard, Xue Li 0001, Mohamed A. Sharaf |
APWeb | 3 |
| 2013 | Scalable diversification of multiple search resultsabstractThe explosion of big data emphasizes the need for scalable data diversification, especially for applications based on web, scientific, and business databases. However, achieving effective diversification in a multi-user environment is a rather challenging task due to the inherent high processing costs of current data diversification techniques. In this paper, we address the concurrent diversification of multiple search results using various approximation techniques that provide orders of magnitude reductions in processing cost, while maintaining comparable quality of diversification as compared to sequential methods. Our extensive experimental evaluation shows the scalability exhibited by our proposed methods under various workload settings. Hina A. Khan, Marina Drosou, Mohamed A. Sharaf |
CIKM | 3 |
| 2013 | Adaptive Query Scheduling in Key-Value Data Stores
Chen Xu 0001, Mohamed A. Sharaf, Minqi Zhou, Aoying Zhou, Xiaofang Zhou 0001 |
DASFAA (1) | 2 |
| 2013 | DoS: an efficient scheme for the diversification of multiple search resultsabstractData diversification provides users with a concise and meaningful view of the results returned by search queries. In addition to taming the information overload, data diversification also provides the benefits of reducing data communication costs as well as enabling data exploration. The explosion of big data emphasizes the need for data diversification in modern data management platforms, especially for applications based on web, scientific, and business databases. Achieving effective diversification, however, is rather a challenging task due to the inherent high processing costs of current data diversification techniques. This challenge is further accentuated in a multi-user environment, in which multiple search queries are to be executed and diversified concurrently. In this paper, we propose the DoS scheme, which addresses the problem of scalable diversification of multiple search results. Our experimental evaluation shows the scalability exhibited by DoS under various workload settings, and the significant benefits it provides compared to sequential methods. Hina A. Khan, Marina Drosou, Mohamed A. Sharaf |
SSDBM | 3 |
| 2012 | Efficient buffer management for piecewise linear representation of multiple data streamsabstractPiecewise Linear Representation (PLR) has been a widely used method for approximating data streams in the form of compact line segments. The buffer-based approach to PLR enables a semi-global approximation which relies on the aggregated processing of batches of streamed data so that to adjust and improve the approximation results. However, one challenge towards applying the buffer-based approach is allocating the necessary memory resources for stream buffering. This challenge is further complicated in a multi-stream environment where multiple data streams are competing for the available memory resources, especially in resource-constrained systems such as sensors and mobile devices. Qing Xie 0002, Jia Zhu 0003, Mohamed A. Sharaf, Xiaofang Zhou 0001, Chaoyi Pang |
CIKM | 3 |
| 2012 | Three-Level Processing of Multiple Aggregate Continuous QueriesabstractAggregate Continuous Queries (ACQs) are both a very popular class of Continuous Queries (CQs) and also have a potentially high execution cost. As such, optimizing the processing of ACQs is imperative for Data Stream Management Systems (DSMSs) to reach their full potential in supporting (critical) monitoring applications. For multiple ACQs that vary in window specifications and pre-aggregation filters, existing multiple ACQs optimization schemes assume a processing model where each ACQ is computed as a final-aggregation of a sub-aggregation. In this paper, we propose a novel processing model for ACQs, called Tri Ops, with the goal of minimizing the repetition of operator execution at the sub-aggregation level. We also propose Tri Weave, a Tri Ops-aware multi-query optimizer. We analytically and experimentally demonstrate the performance gains of our proposed schemes which shows their superiority over alternative schemes. Finally, we generalize Tri Weave to incorporate the classical subsumption-based multi-query optimization techniques. Shenoda Guirguis, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis |
ICDE | 2 |
| 2012 | RFID Mutual Authentication and Secret Update Protocol for Low-Cost TagsabstractLow-cost RFID will dominate the industry as a replacement of the barcode tags. RFID tags had been designed for the purpose of automatic identification and tracking. Therefore, RFID tags could violate their owners' privacy and security. Hence, it becomes a necessity to come up with an RFID protocol that meets security (e.g. mutual authentication) and privacy goals. Recently, the Journal of Computer Communications published a paper by Song and Mitchell (2011) where the authors proposed a mutual authentication protocol for RFID system. This protocol has fundamental shortcomings that can be taken advantage by an adept adversary. We shed light on these flaws and later on we modified SM's scheme to fix these vulnerabilities. Mohamed A. Sharaf |
TrustCom | 1 |
| 2012 | WebPut: Efficient Web-Based Data Imputation
Zhixu Li, Mohamed A. Sharaf, Laurianne Sitbon, Shazia Sadiq, Marta Indulska, Xiaofang Zhou 0001 |
WISE | 2 |
| 2012 | On the Prediction of Re-tweeting Activities in Social Networks - A Report on WISE 2012 Challenge
Sayan Unankard, Ling Chen 0004, Sen Wang 0001, Zi Huang, Mohamed A. Sharaf, Xue Li 0001 |
WISE | 6 |
| 2011 | Optimized processing of multiple aggregate continuous queriesabstractData Streams Management Systems are designed to support monitoring applications, which require the processing of hundreds of Aggregate Continuous Queries (ACQs). These ACQs typically have different time granularities, with possibly different selection predicates and group-by attributes. In order to achieve scalability in the presence of heavy workloads, in this paper, we introduce the concept of 'Weaveability' as an indicator of the potential gains of sharing the processing of ACQs. We then propose Weave Share, a cost-based optimizer that exploits weaveability to optimize the shared processing of ACQs. Our experimental analysis shows that Weave Share outperforms the alternative sharing schemes generating up to four orders of magnitude better quality plans. Finally, we describe a practical implementation of the Weave Share optimizer. Shenoda Guirguis, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis |
CIKM | 2 |
| 2011 | Optimizing the Energy Consumption of Continuous Query Processing with Mobile ClientsabstractComplex event detection over data streams has become ubiquitous through the widespread use of sensors, wireless connectivity and the wide variety of end-user mobile devices. Typically, such event detection is carried out by a data stream management system executing continuous queries (CQs), registered by the users. In this paper, we consider the situation where the results of the CQs, which are in the form of individual data streams, are disseminated to the users' hand-held, battery-operated devices over a shared broadcast medium. In order to reduce the overall energy consumption of the mobile devices, we propose Bose*, a power-aware query operator placement algorithm that determines which part of a CQ plan should be executed at the data stream management system and which part should be executed at the mobile device. Bose*'s effectiveness in reducing energy consumption, as well as response time under specific conditions, is evaluated using simulation, driven by parameters measured on real mobile devices. Panayiotis Neophytou, Jesse Szwedko, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis |
Mobile Data Management (1) | 3 |
| 2011 | Visualization of Energy Consumption of Continuous Query Processing with Mobile ClientsabstractComplex event detection over data streams has become ubiquitous through the widespread use of sensors, wireless connectivity and the wide variety of end-user mobile devices. Typically, event detection is carried out by a central server executing continuous queries. In this demonstration, we focus on the case where users with mobile devices submit continuous queries (for event detection) to a data stream management server which disseminates the results to the users over a shared broadcast medium. In order to minimize the overall energy consumption of the mobile devices (clients), we have proposed operator placement algorithms that split the processing of each continuous query between the centralized server and the requesting mobile clients, thus trading off energy consumption for communication energy consumption for computation. Specifically, in this demonstration, we present an interactive graphical interface to the inner workings of our three proposed operator placement algorithms, whereby attendees are able to investigate various query plans and the decisions that the algorithms make, as well as visualize the results of these algorithms in terms of client power consumption and response time. Besides being able to step through an algorithm's execution as it considers various operator placement decisions, attendees are able to experiment with different scenarios by customizing the parameters of the query workloads (e.g., changing the selectivities and projectivities of the operators) or the client's profile (e.g., power consumed per unit of time of processing) and examine the impact. Jesse Szwedko, Panayiotis Neophytou, Panos K. Chrysanthis, Alexandros Labrinidis, Mohamed A. Sharaf |
Mobile Data Management (1) | 5 |
| 2009 | Adaptive Scheduling of Web TransactionsabstractIn highly interactive dynamic Web database systems, user satisfaction determines their success. In such systems, user requested web pages are dynamically created by executing a number of database queries or Web transactions. In this paper, we model the interrelated transactions generating a web page asworkflowsand quantify the user satisfaction by associating dynamic Web pages withsoft-deadlines. Further, we model the importance of transactions in generating a page by associating different weights to transactions. Using this framework, system success is measured in terms of minimizing the deviation from the deadline (i.e., tardiness) and also minimizing the weighted such deviation (i.e., weighted tardiness). In order to efficiently support the materialization of dynamic Web pages, we proposeASETS*, which is a parameter-free adaptive scheduling algorithm that automatically adapts to, not only system load, but also transactions' characteristics (i.e., interdependencies, deadlines and weights).ASETS* prioritizes the execution of transactions with the objective of minimizing weighted tardiness. It is also capable of balancing the tradeoff between optimizing average- and worst-case performance when needed. The performance advantages ofASETS* are experimentally demonstrated. Shenoda Guirguis, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Kirk Pruhs |
ICDE | 2 |
| 2009 | SLA-Aware Adaptive On-demand Data Broadcasting in Wireless EnvironmentsabstractIn mobile and wireless networks, data broadcasting for popular data items enables the efficient utilization of the limited wireless bandwidth. However, efficient data scheduling schemes are needed to fully exploit the benefits of data broadcasting. This motivated the proposal of several broadcast scheduling policies, which have mostly focused on either minimizing response time, or drop rate when requests are associated with hard deadlines. The inherent inaccuracy of hard deadlines in a dynamic mobile environment motivated us to use Service Level Agreements (SLAs) where a user specifies the utility of data as a function of its arrival time. Moreover, SLAs provide the mobile user with an already familiar quality of service specification from wired environments. Hence, in this paper, we propose SAAB which is an SLA-aware adaptive data broadcast scheduling policy for maximizing the system utility under SLA-based performance measures. To achieve this goal, SAAB considers both the characteristics of disseminated data objects as well as the SLAs associated with them. Additionally, SAAB automatically adjusts to the system workload conditions which enables it to constantly outperform existing on-demand broadcast scheduling policies. Adrian Daniel Popescu, Mohamed A. Sharaf, Cristiana Amza |
Mobile Data Management | 2 |
| 2009 | Optimizing i/o-intensive transactions in highly interactive applicationsabstractThe performance provided by an interactive online database system is typically measured in terms of meeting certain pre-specified Service Level Agreements (SLAs), with expected transaction latency being the most commonly used type of SLA. This form of SLA acts as a soft deadline for each transaction, and user satisfaction can be measured in terms of minimizing tardiness, that is, the deviation from SLA. This objective is further complicated for I/O-intensive transactions, where the storage system becomes the performance bottleneck. Moreover, common I/O scheduling policies employed by the Operating System with a goal of improving I/O throughput or average latency may run counter to optimizing per-transaction performance since the Operating System is typically oblivious to the application high-level SLA specifications. In this paper, we propose a new SLA-aware policy for scheduling I/O requests of database transactions. Our proposed policy synergistically combines novel deadline-aware scheduling policies for database transactions with features of Operating System scheduling policies designed for improving I/O throughput. This enables our proposed policy to dynamically adapt to workload and consistently provide the best performance. Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Cristiana Amza |
SIGMOD Conference | 1 |
| 2008 | Scheduling continuous queries in data stream management systemsabstractRecently, several policies have been proposed for scheduling multiple Continuous Queries (CQs) in a Data Stream Management System (DSMS). The decision on which policy to use plays an important role in shaping the percieved online performance provided by the DSMS. In this tutorial, we provide an overview of different policies employed by current CQ schedulers and the performance goals optimized by these policies. Further, we discuss the salient properties of CQs conisdered by current policies as well as the efficent implementation of such policies into CQ schedulers. Finally, we present future research directions and open problems in CQ scheduling. Mohamed A. Sharaf, Alexandros Labrinidis, Panos K. Chrysanthis |
Proc. VLDB Endow. | 1 |
| 2008 | Dynamic partitioning of the cache hierarchy in shared data centersabstractDue to the imperative need to reduce the management costs of large data centers, operators multiplex several concurrent database applications on a server farm connected to shared network attached storage. Determining and enforcing per-application resource quotas in the resulting cache hierarchy, on the fly, poses a complex resource allocation problem spanning the database server and the storage server tiers. This problem is further complicated by the need to provide strict Quality of Service (QoS) guarantees to hosted applications. In this paper, we design and implement a novel coordinated partitioning technique of the database buffer pool and storage cache between applications for any given cache replacement policy and per-application access pattern. We use statistical regression to dynamically determine the mapping between cache quota settings and the resulting per-application QoS. A resource controller embedded within the database engine actuates the partitioning of the two-level cache, converging towards the configuration with maximum application utility, expressed as the service provider revenue in that configuration, based on a set of latency sample points. Our experimental evaluation, using the MySQL database engine, a server farm with consolidated storage, and two e-commerce benchmarks, shows the effectiveness of our technique in enforcing application QoS, as well as maximizing the revenue of the service provider in shared server farms. Gokul Soundararajan, Jin Chen 0006, Mohamed A. Sharaf, Cristiana Amza |
Proc. VLDB Endow. | 3 |
| 2008 | Algorithms and metrics for processing multiple heterogeneous continuous queriesabstractThe emergence of monitoring applications has precipitated the need for Data Stream Management Systems (DSMSs), which constantly monitor incoming data feeds (through registered continuous queries), in order to detect events of interest. In this article, we examine the problem of how to schedule multiple Continuous Queries (CQs) in a DSMS to optimize different Quality of Service (QoS) metrics. We show that, unlike traditional online systems, scheduling policies in DSMSs that optimize for average response time will be different from policies that optimize for average slowdown, which is a more appropriate metric to use in the presence of a heterogeneous workload. Towards this, we propose policies to optimize for the average-case performance for both metrics. Additionally, we propose a hybrid scheduling policy that strikes a fine balance between performance and fairness, by looking at both the average- and worst-case performance, for both metrics. We also show how our policies can be adaptive enough to handle the inherent dynamic nature of monitoring applications. Furthermore, we discuss how our policies can be efficiently implemented and extended to exploit sharing in optimized multi-query plans and multi-stream CQs. Finally, we experimentally show using real data that our policies consistently outperform currently used ones. Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Kirk Pruhs |
ACM Trans. Database Syst. | 1 |
| 2006 | Efficient Scheduling of Heterogeneous Continuous Queries
Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Kirk Pruhs |
VLDB | 1 |
| 2005 | Preemptive rate-based operator scheduling in a data stream management systemabstractSummary form only given. Data stream management systems are being developed to process continuous queries over multiple data streams. These continuous queries are typically used for monitoring purposes where the detection of an event might trigger a sequence of actions or the execution of a set of specified tasks. Such events are identified by tuples produced by a query and hence, it is important to produce the available portions of a query result as early as possible. A core element for improving the interactive performance of a continuous query is the operator scheduler. An operator scheduler is particularly important when the processing requirements and the productivity of different streams are highly skewed. The need for an operator scheduler becomes even more crucial when tuples from different streams arrive asynchronously. To meet these needs, we are proposing a preemptive rate-based scheduling policy that handles the asynchronous nature of tuple arrival and the heterogeneity in the query plan. Experimental results show the significant improvements provided by our proposed policy. Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis |
AICCSA | 1 |
| 2005 | Freshness-Aware Scheduling of Continuous Queries in the Dynamic Web
Mohamed A. Sharaf, Alexandros Labrinidis, Panos K. Chrysanthis, Kirk Pruhs |
WebDB | 1 |
| 2004 | On-Demand Data Broadcasting for Mobile Decision Making
Mohamed A. Sharaf, Panos K. Chrysanthis |
Mob. Networks Appl. | 1 |
| 2004 | Balancing energy efficiency and quality of aggregate data in sensor networks
Mohamed A. Sharaf, Jonathan Beaver, Alexandros Labrinidis, Panos K. Chrysanthis |
VLDB J. | 1 |
| 2003 | An Optimized Multicast-based Data Dissemination MiddlewareabstractA major problem on the Internet is the scalable dissemination of information. This problem is particularly acute exactly at the time when the scalability of data delivery is most important. One proposed solution to this scalability problem is to use multicast communication. However, allowing multicast communication introduces many nontrivial data management problems, such as caching, consistency, and scheduling. We have built a middleware that unifies and extends state-of-the-art data management methods and algorithms into one software distribution. Its flexible and extensible architecture is built from individual components that can be selected or replaced depending on the underlying multicast transport mechanism or on the application needs. Particular care has gone into the design of the algorithms to optimize the user-perceived level of service. We demonstrate our middleware within the context of the RODS application. Wenhui Zhang 0002, Vincenzo Liberatore, Vince Penkrot, Jonathan Beaver, Mohamed A. Sharaf, Siddhartha Roychowdhury, Panos K. Chrysanthis, Kirk Pruhs |
ICDE | 6 |
| 2003 | Efficient Dissemination of Aggregate Data over the Wireless Web
Mohamed A. Sharaf, Yannis Sismanis, Alexandros Labrinidis, Panos K. Chrysanthis, Nick Roussopoulos |
WebDB | 1 |
| 2002 | Semantic-based delivery of OLAP summary tables in wireless environmentsabstractWith the rapid growth in mobile and wireless technologies and the availability, pervasiveness and cost effectiveness of wireless networks, mobile computers are quickly becoming the normal front-end devices for accessing enterprise data. In this paper, we are addressing the issue of efficient delivery of business decision support data in the form of summary tables to mobile clients equipped with OLAP front-end tools. Towards this, we propose a new on-demand scheduling algorithm, called SBS, that exploits both the derivation semantics among OLAP summary tables and the mobile clients' capabilities of executing simple SQL queries. It maximizes the aggregated data sharing between clients and reduces the broadcast length compared to the already existing techniques. The degree of aggregation can be tuned to control the tradeoff between access time and energy consumption. Further, the proposed scheme adapts well to different request rates, access patterns and data distributions. The algorithm effectiveness with respect to access time and power consumption is evaluated using simulation. Mohamed A. Sharaf, Panos K. Chrysanthis |
CIKM | 1 |