Mohamed A. Sharaf

dblp:s/MohamedASharaf · DBLP profile ↗
← Back
54ranked-venue papers
12as first author
4since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 44 · 9 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 1 first-authorComputer networks · 2 · 1 first-authorSecurity and privacy · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
13 papers
Information retrieval · 51% Query processing and optimization · 16% Data stream processing · 9%
Computer architecture, parallel and distributed computing, and storage systems
7 papers
Storage systems · 46% Cloud and datacenter computing · 32% Distributed systems · 10%
Computer graphics and multimedia
1 paper
Visualization and visual analytics · 50% Computational photography and imaging · 50%

Topics — the 30 heaviest of 39, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
query formulation
0.622020
Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020
CoRE: A Context-Aware RelationExtraction Method for Relation Completion · IEEE Trans. Knowl. Data Eng. 2014
Information retrieval › query formulation › query generation
boolean query generation
0.412020
Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020
Information retrieval › ranking
learning to rank
0.412020
Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020
Information retrieval
retrieval models
0.412020
Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020
Information retrieval
search engines
0.412020
Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews · WWW 2020
Data integration and cleaning › missing data
missing value imputation
0.412019
WebPut: A Web-Aided Data Imputation System for the General Type of Missing String Attribute Values · ICDE 2019
Computational photography and imaging
view recommendation
0.312018
Efficient Recommendation of Aggregate Data Visualizations · IEEE Trans. Knowl. Data Eng. 2018
Visualization and visual analytics
visualization recommendation
0.312018
Efficient Recommendation of Aggregate Data Visualizations · IEEE Trans. Knowl. Data Eng. 2018
Recommender systems › domain-specific recommendation
visualization recommendation
0.212016
MuVE: Efficient Multi-Objective View Recommendation for Visual Data Exploration · ICDE 2016
Data stream processing › continuous query processing
continuous query scheduling
0.232008
Algorithms and metrics for processing multiple heterogeneous continuous queries · ACM Trans. Database Syst. 2008
Scheduling continuous queries in data stream management systems · Proc. VLDB Endow. 2008
Efficient Scheduling of Heterogeneous Continuous Queries · VLDB 2006
Query processing and optimization
multi-query optimization
0.222012
Three-Level Processing of Multiple Aggregate Continuous Queries · ICDE 2012
Algorithms and metrics for processing multiple heterogeneous continuous queries · ACM Trans. Database Syst. 2008
Information retrieval
search result diversification
0.212015
Progressive diversification for column-based data exploration platforms · ICDE 2015
Natural language and speech › Information extraction and text analysis
relation extraction
0.212014
CoRE: A Context-Aware RelationExtraction Method for Relation Completion · IEEE Trans. Knowl. Data Eng. 2014
Data mining › dimensionality reduction
feature selection
0.212014
Mining Personal Health Index from Annual Geriatric Medical Examinations · ICDM 2014
Cloud and datacenter computing
cluster resource management and scheduling
0.212014
AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014
Storage systems
key-value storage
0.212014
AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014
Storage systems › key-value storage
NoSQL database
0.212014
AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014
Data stream processing › continuous query processing
continuous aggregate query
0.112012
Three-Level Processing of Multiple Aggregate Continuous Queries · ICDE 2012
Transaction processing and concurrency control
transaction scheduling
0.112009
Optimizing i/o-intensive transactions in highly interactive applications · SIGMOD Conference 2009
Storage systems
i/o scheduling
0.112009
Optimizing i/o-intensive transactions in highly interactive applications · SIGMOD Conference 2009
Indexing and storage engines
buffer management
0.112008
Dynamic partitioning of the cache hierarchy in shared data centers · Proc. VLDB Endow. 2008
Query processing and optimization › multi-query optimization
query plan sharing
0.112008
Algorithms and metrics for processing multiple heterogeneous continuous queries · ACM Trans. Database Syst. 2008
Data stream processing
stream processing systems
0.112008
Scheduling continuous queries in data stream management systems · Proc. VLDB Endow. 2008
Memory systems › cache management
cache partitioning
0.112008
Dynamic partitioning of the cache hierarchy in shared data centers · Proc. VLDB Endow. 2008
Cloud and datacenter computing
resource management
0.112008
Dynamic partitioning of the cache hierarchy in shared data centers · Proc. VLDB Endow. 2008
Data mining › exploratory data analysis
visual exploration
0.112016
MuVE: Efficient Multi-Objective View Recommendation for Visual Data Exploration · ICDE 2016
Distributed systems
replication
0.112014
AQUAS: A quality-aware scheduler for NoSQL data stores · ICDE 2014
Internet of things and sensor networks › wireless sensor network
in-network aggregation
0.012004
Balancing energy efficiency and quality of aggregate data in sensor networks · VLDB J. 2004
Cellular and mobile networks
scalable information dissemination
0.012003
An Optimized Multicast-based Data Dissemination Middleware · ICDE 2003
Distributed systems › distributed communication
data dissemination
0.012003
An Optimized Multicast-based Data Dissemination Middleware · ICDE 2003

Methods — techniques the papers use, named apart from their topics

query variation sampling · 0.4learning to rank · 0.4crowdsourcing · 0.4pattern-based extraction · 0.4model selection · 0.4data preprocessing · 0.4context term learning · 0.4pruning · 0.3multi-objective optimization · 0.3multi-objective utility function · 0.2incremental pruning · 0.2partial distance computation · 0.2i/o throughput optimization · 0.2deadline-aware scheduling · 0.2quality-of-service scheduling · 0.2quality-of-data management · 0.2tardiness minimization · 0.1adaptive scheduling · 0.1
YearPublicationVenuePosition
2025 From Incomplete Data to Accurate Aggregation: Exploring the Accuracy-Efficiency Trade-off of Data Imputation Methods
abstract
Missing data remains a critical challenge in data analytics, especially for aggregation-based descriptive tasks. In addition to evaluating the accuracy of data imputation methods at the cell-level, several works have studied their impact on downstream predictive analytics. However, the interplay between data imputation and data aggregation remains largely underexplored. To address that limitation, in this work, we evaluate the impact of data imputation on the accuracy of aggregate queries. Our evaluation is conducted on several real datasets, measuring the accuracy at both the cell-level and the aggregatelevel, as well as runtime cost. Our results show clear trade-offs between accuracy and efficiency, providing guidance for selecting imputation strategies in aggregation-based analytical pipelines.
Linda Mohammed, Heba Helal, Mohamed A. Sharaf
AICCSA3
2023 Strategies for Optimizing Time Series Visual Data Exploration
abstract
Dealing with high-dimensional time series data makes the process of recommending visualizations with "interesting" insights difficult. The challenge originates from finding a way to obtain the recommended visualizations efficiently without compromising their quality. Identifying such visualizations manually is considered a labor-intensive and time-consuming process. In response, this paper introduces different techniques designed to optimize the automated recommendation process. These techniques are entirely based on the concept of computation sharing and pruning. Furthermore, we provide a glimpse into our future research works in PhD thesis. The objective is to broaden the scope of our current work and enhance the generality of our problem statement.
Heba Helal, Mohamed A. Sharaf
AICCSA2
2023 TiVEx: Optimized Processing for Time Series Visual Exploration
abstract
To facilitate fast-visual data analysis, there is a need for recommending top-k views with "interesting" insights automatically. However, working with high-dimensional time series data makes the process of view recommendations difficult. The primary obstacle lies in finding an automatic way to generate views with less processing time (efficiency) while still closely aligning with the ground truth (effectiveness). In this paper, we propose TiVEx (Time Series Visual Exploration), a technique to address this challenge. TiVEx aims to achieve a balance between efficiency and effectiveness in generating view recommendations. Through extensive experiments, we demonstrate significant cost savings achieved by TiVEx, indicating its efficiency. Furthermore, our analysis delves into the exploration of striking the right balance between efficiency and effectiveness.
Heba Helal, Mohamed A. Sharaf, Mohammad M. Masud 0001, Panos K. Chrysanthis
AICCSA2
2022 CovidLens: Visually Understanding the Covid-19 Indicators through the Lens of Mobility Data
abstract
Since the onset of the Covid-19 pandemic, an over-whelming amount of related data has been released. In an attempt to gain insights from that data, multiple public data visualization dashboards have been deployed. Differently from such dashboards, which mainly support basic data filtering and visualization of separate datasets, in this work, we propose CovidLens, which: 1) integrates various Covid-19 indicators and is centred around the Google Community Mobility Report dataset, 2) supports similarity search for finding similar and correlated patterns and trends across the integrated datasets, and 3) automatically recommends insightful visualizations that unlocks valuable insights into the pandemic effects. To that end, we will be presenting the employed dataset, together with the design, implementation, and multiple usage scenarios of our proposed CovidLens.
Mohamed A. Sharaf, Xiaozhong Zhang, Panos K. Chrysanthis, Wadima Alsaedi, Maitha Alkalbani, Heba Helal, Alyazia Aldhaheri
MDM1
2020 Quality Matters: Understanding the Impact of Incomplete Data on Visualization Recommendation
Rischan Mafrur, Mohamed A. Sharaf, Guido Zuccon
DEXA (1)2
2020 Sampling Query Variations for Learning to Rank to Improve Automatic Boolean Query Generation in Systematic Reviews
abstract
Searching medical literature for synthesis in a systematic review is a complex and labour intensive task. In this context, expert searchers construct lengthy Boolean queries. The universe of possible query variations can be massive: a single query can be composed of hundreds of field-restricted search terms/phrases or ontological concepts, each grouped by a logical operator nested to depths of sometimes five or more levels deep. With the many choices about how to construct a query, it is difficult to both formulate and recognise effective queries. To address this challenge, automatic methods have recently been explored for generating and selecting effective Boolean query variations for systematic reviews. The limiting factor of these methods is that it is computationally infeasible to process all query variations for training the methods. To overcome this, we propose novel query variation sampling methods for training Learning to Rank models to rank queries. Our results show that query sampling methods do directly impact the ability of a Learning to Rank model to effectively identify good query variations. Thus, selecting appropriate query sampling methods is a key problem for the automatic reformulation of effective Boolean queries for systematic review literature search. We find that the best sampling strategies are those which balance the diversity of queries with the quantity of queries.
Harrisen Scells, Guido Zuccon, Mohamed A. Sharaf, Bevan Koopman
WWW3
2020 Serendipity-based Points-of-Interest Navigation
abstract
Traditional venue and tour recommendation systems do not necessarily provide a diverse set of recommendations and leave little room for serendipity . In this article, we design MPG, a Mobile Personal Guide that recommends: (i) a set of diverse yet surprisingly interesting venues that are aligned to user preferences and (ii) a set of routes, constructed from the recommended venues. We also introduce EPUI, an Experimental Platform for Urban Informatics. Our comparison with the state-of-the-art schemes indicates that MPG is capable of providing high-quality venues and route recommendations while incorporating seamlessly both the notion of diversity and that of serendipity.
Xiaoyu Ge, Panos K. Chrysanthis, Konstantinos Pelechrinis, Demetris Zeinalipour, Mohamed A. Sharaf
ACM Trans. Internet Techn.5
2019 WebPut: A Web-Aided Data Imputation System for the General Type of Missing String Attribute Values
abstract
In this demonstration, we present an end-to-end web-aided data imputation prototype system named WebPut. WebPut consults the Web for imputing the missing values in a local database when the traditional inferring-based imputation method has difficulties in getting the right answers. Specifically, WebPut investigates the interaction between the local inferring-based imputation methods and the web-based retrieving methods and shows that retrieving a small number of selected missing values can greatly improve the imputation recall of the inferring-based methods. Besides, WebPut also incorporates a crowd intervention component that can get advice from humans in case that the web-based imputation methods may have difficulties in making the right decisions. We demonstrate, step by step, how WebPut fills an incomplete table with each of its components.
Shuangli Shan, Zhixu Li, Qiang Yang 0015, Jia Zhu 0003, Mohamed A. Sharaf, Xiaofang Zhou 0001
ICDE6
2018 DiVE: Diversifying View Recommendation for Visual Data Exploration
abstract
To support effective data exploration, there has been a growing interest in developing solutions that can automatically recommend data visualizations that reveal interesting and useful data-driven insights. In such solutions, a large number of possible data visualization views are generated and ranked according to some metric of importance (e.g., a deviation-based metric), then the top-k most important views are recommended. However, one drawback of that approach is that it often recommends similar views, leaving the data analyst with a limited amount of gained insights. To address that limitation, in this work we posit that employing diversification techniques in the process of view recommendation allows eliminating that redundancy and provides a good and concise coverage of the possible insights to be discovered. To that end, we propose a hybrid objective utility function, which captures both the importance, as well as the diversity of the insights revealed by the recommended views. While in principle, traditional diversification methods (e.g., Greedy Construction) provide plausible solutions under our proposed utility function, they suffer from a significantly high query processing cost. In particular, directly applying such methods leads to a "process-first-diversify-next" approach, in which all possible data visualization are generated first via executing a large number of aggregate queries. To address that challenge, we propose an integrated scheme called DiVE, which efficiently selects the top-k recommended view based on our hybrid utility function. DiVE leverages the properties of both the importance and diversity metrics to prune a large number of query executions without compromising the quality of recommendations. Our experimental evaluation on real datasets shows the performance gains provided by DiVE.
Rischan Mafrur, Mohamed A. Sharaf, Hina A. Khan
CIKM2
2018 Efficient Recommendation of Aggregate Data Visualizations
abstract
Data visualization is a common and effective technique for data exploration. However, for complex data, it is infeasible for an analyst to manually generate and browse all possible visualizations for insights. This observation motivated the need for automated solutions that can effectively recommend such visualizations. The main idea underlying those solutions is to evaluate the utility of all possible visualizations and then recommend the top-k visualizations. This process incurs high data processing cost, that is further aggravated by the presence of numerical dimensional attributes. To address that challenge, we propose novel view recommendation schemes, which incorporate a hybrid multi-objective utility function that captures the impact of numerical dimension attributes. Our first scheme, Multi-Objective View Recommendation for Data Exploration (MuVE), adopts an incremental evaluation of our multi-objective utility function, which allows pruning of a large number of low-utility views and avoids unnecessary objective evaluations. Our second scheme, upper MuVE (uMuVE), further improves the pruning power by setting the upper bounds on the utility of views and allowing interleaved processing of views, at the expense of increased memory usage. Finally, our third scheme, Memory-aware uMuVE (MuMuVE), provides pruning power close to that of uMuVE, while keeping memory usage within a pre-specified limit.
Humaira Ehsan, Mohamed A. Sharaf, Panos K. Chrysanthis
IEEE Trans. Knowl. Data Eng.2
2017 In Search for Relevant, Diverse and Crowd-screen Points of Interests
abstract
In this demo we present a prototype of an experimental platform for evaluating item recommendation algorithms. The application domain for our system is that of digital city guides. Our prototype implementation allows the user to explore different algorithms and compare their output. Among the algorithms implemented is MPG, which aims at providing a diverse set of recommendations better aligned with user preferences. MPG takes into consideration the user preferences (e.g., reach willing to cover, types of venues interested in exploring etc.), the popularity of the establishments as well as their distance from the current location of the user by combining them into a single composite score. We provide a web interface, which outputs on a map the recommended locations along with metadata (e.g., type and name of location, relevance and diversity scores, etc.). It also illustrates the potential of the Preferential Diversity approach on which MPG is based.
Xiaoyu Ge, Samanvoy Panati, Konstantinos Pelechrinis, Panos K. Chrysanthis, Mohamed A. Sharaf
EDBT5
2017 Model-Based Diversification for Sequential Exploratory Queries
abstract
Today, data exploration platforms are widely used to assist users in locating interesting objects within large volumes of scientific and business data. In those platforms, users try to make sense of the underlying data space by iteratively posing numerous queries over large databases. While diversification of query results, like other data summarization techniques, provides users with quick insights into the huge query answer space, it adds additional complexity to an already computationally expensive data exploration task. To address this challenge, in this paper we propose a diversification scheme that targets the problem of efficiently diversifying the results of multiple queries within and across different data exploratory sessions. Our proposed scheme relies on a model-based diversification method and an ordered cache. In particular, we employ an adaptive regression model to estimate the diversity of a diverse subset. Such estimation of diversity value allows us to select diverse results without scanning all the query results. In order to further expedite the diversification process, we propose an order-based caching scheme to leverage the overlap between sequence of data exploration queries. Our extensive experimental evaluation on both synthetic and real data sets shows the significant benefits provided by our scheme as compared to the existing methods.
Hina A. Khan, Mohamed A. Sharaf
Data Sci. Eng.2
2017 Efficient schemes for similarity-aware refinement of aggregation queries
Abdullah M. Albarrak, Mohamed A. Sharaf
World Wide Web2
2016 REQUEST: A scalable framework for interactive construction of exploratory queries
abstract
Exploration over large datasets is a key first step in data analysis, as users may be unfamiliar with the underlying database schema and unable to construct precise queries that represent their interests. Such data exploration task usually involves executing numerous ad-hoc queries, which requires a considerable amount of time and human effort. In this paper, we present REQUEST, a novel framework that is designed to minimize the human effort and enable both effective and efficient data exploration. REQUEST supports the query-from-examples style of data exploration by integrating two key components: 1) Data Reduction, and 2) Query Selection. As instances of the REQUEST framework, we propose several highly scalable schemes, which employ active learning techniques and provide different levels of efficiency and effectiveness as guided by the user's preferences. Our results, on real-world datasets from Sloan Digital Sky Survey, show that our schemes on average require 1-2 orders of magnitude fewer feedback questions than the random baseline, and 3-16× fewer questions than the state-of-the-art, while maintaining interactive response time. Moreover, our schemes are able to construct, with high accuracy, queries that are often undetectable by current techniques.
Xiaoyu Ge, Yanbing Xue, Mohamed A. Sharaf, Panos K. Chrysanthis
IEEE BigData4
2016 MuVE: Efficient Multi-Objective View Recommendation for Visual Data Exploration
abstract
To support effective data exploration, there is a well-recognized need for solutions that can automatically recommend interesting visualizations, which reveal useful insights into the analyzed data. However, such visualizations come at the expense of high data processing costs, where a large number of views are generated to evaluate their usefulness. Those costs are further escalated in the presence of numerical dimensional attributes, due to the potentially large number of possible binning aggregations, which lead to a drastic increase in the number of possible visualizations. To address that challenge, in this paper we propose the MuVE scheme for Multi-Objective View Recommendation for Visual Data Exploration. MuVE introduces a hybrid multi-objective utility function, which captures the impact of binning on the utility of visualizations. Consequently, novel algorithms are proposed for the efficient recommendation of data visualizations that are based on numerical dimensions. The main idea underlying MuVE is to incrementally and progressively assess the different benefits provided by a visualization, which allows an early pruning of a large number of unnecessary operations. Our extensive experimental results show the significant gains provided by our proposed scheme.
Humaira Ehsan, Mohamed A. Sharaf, Panos K. Chrysanthis
ICDE2
2015 Progressive diversification for column-based data exploration platforms
abstract
In Data Exploration platforms, diversification has become an essential method for extracting representative data, which provide users with a concise and meaningful view of the results to their queries. However, the benefits of diversification are achieved at the expense of an additional cost for the post-processing of query results. For high dimensional large result sets, the cost of diversification is further escalated due to massive distance computations required to evaluate the similarity between results. To address that challenge, in this paper we propose the Progressive Data Diversification (pDiverse) scheme. The main idea underlying pDiverse is to utilize partial distance computation to reduce the amount of processed data. Our extensive experimental results on both synthetic and real data sets show that our proposed scheme outperforms existing diversification methods in terms of both I/O and CPU costs.
Hina A. Khan, Mohamed A. Sharaf
ICDE2
2015 Emerging event detection in social networks with location sensitivity
Sayan Unankard, Xue Li 0001, Mohamed A. Sharaf
World Wide Web3
2014 SAQR: An Efficient Scheme for Similarity-Aware Query Refinement
Abdullah M. Albarrak, Mohamed A. Sharaf, Xiaofang Zhou 0001
DASFAA (1)2
2014 AQUAS: A quality-aware scheduler for NoSQL data stores
abstract
NoSQL key-value data stores provide an attractive solution for big data management. With the help of data partitioning and replication, those data stores achieve higher levels of availability, scalability and reliability. Such design choices typically exhibit a tradeoff in which data freshness is sacrificed in favor of reduced access latency. At the replica-level, this tradeoff is primarily shaped by the resource allocation strategies deployed for managing the processing of user queries and replica updates. In this demonstration, we showcase AQUAS: a quality-aware scheduler for Cassandra, which allows application developers to specify requirements on quality of service (QoS) and quality of data (QoD). AQUAS efficiently allocates the available replica resources to execute the incoming read/write tasks so that to minimize the penalties incurred by violating those requirements. We demonstrate AQUAS based on our implementation of a microblogging system.
Chen Xu 0001, Mohamed A. Sharaf, Minqi Zhou, Aoying Zhou
ICDE3
2014 Mining Personal Health Index from Annual Geriatric Medical Examinations
abstract
People take regular medical examinations mostly not for discovering diseases but for having a peace of mind regarding their health status. Therefore, it is important to give them an overall feedback with respect to all the health indicators that have been ranked against the whole population. In this paper, we propose a framework of mining Personal Health Index (PHI) from a large and comprehensive geriatric medical examination (GME) dataset. We define PHI as an overall score of personal health status based on a complement probability of health risks. The health risks are calculated using the information from the cause of death (COD) dataset that is linked to the GME dataset. Especially, the highest health risk is revealed in the cases of people who had been taking GME for some years and then passed away for medical reasons. The proposed framework consists of methods in data pre-processing, feature extraction and selection, and model selection. The effectiveness of the proposed framework is validated by a set of comprehensive experiments based on the records of 102,258 participants. As the first of this kind, our work provides a baseline for further research.
Ling Chen 0004, Xue Li 0001, Sen Wang 0001, Hsiao-Yun Hu, Nicole Huang, Quan Z. Sheng, Mohamed A. Sharaf
ICDM7
2014 ORange: Objective-Aware Range Query Refinement
abstract
In this demo paper we present Orange, a system prototype for objective-aware range query refinement. Orange essentially refines a range query to meet a pre-specified cardinality constraint while taking into account the (dis)similarity between the initial query and its corresponding refined version. To achieve this goal, Orange employes the novel scheme SAQR for efficient similarity-aware query refinement. The main idea underlying SAQR is to utilize the pre-defined constraints on cardinality and similarity in order to bound the search space and quickly find a refined query, which meets the user's expectations. We showcase Orange in a web-based application which aims to guide planners in allocating service zones for police patrol units using real and historical dataset of crime incidents.
Abdullah M. Albarrak, Tatiana Noboa, Hina A. Khan, Mohamed A. Sharaf, Xiaofang Zhou 0001, Shazia Sadiq
MDM (1)4
2014 Efficient Retrieval of Top-K Most Similar Users from Travel Smart Card Data
abstract
Understanding the dynamics of human daily mobility patterns is essential for the management and planning of urban facilities and services. Travel smart cards, which record users' public transporting histories, capture rich information of users' mobility pattern. This provides the opportunity to discover valuable knowledge from these transaction records. In recent years, research on measuring user similarity for behavior analysis has attracted a lot of attention in applications such as recommendation systems, crowd behavior analysis applications, and numerous data mining tasks. In this paper, our goal is to estimate the similarity between users' travel patterns according to their travel smart card data. The core of our proposal is a novel user similarity measurement, namely, Travel Spatial-Temporal Similarity (TST), which measures the spatial range and temporal similarity between users. Moreover, we also propose a hybrid index structure, which integrates inverted files and cluster-based partitioning, to allow for efficient retrieval of the top-K most similar users. Through experimental evaluation, our proposed approach is shown to deliver scalable performance.
Bolong Zheng, Kai Zheng 0001, Mohamed A. Sharaf, Xiaofang Zhou 0001, Shazia Sadiq
MDM (1)3
2014 DivIDE: efficient diversification for interactive data exploration
abstract
Today, Interactive Data Exploration (IDE) has become a main constituent of many discovery-oriented applications, in which users repeatedly submit exploratory queries to identify interesting subspaces in large data sets. Returning relevant yet diverse results to such queries provides users with quick insights into a rather large data space. Meanwhile, search results diversification adds additional cost to an already computationally expensive exploration process. To address this challenge, in this paper, we propose a novel diversification scheme called DivIDE, which targets the problem of efficiently diversifying the results of queries posed during data exploration sessions. In particular, our scheme exploits the properties of data diversification functions while leveraging the natural overlap occurring between the results of different queries so that to provide significant reductions in processing costs. Our extensive experimental evaluation on both synthetic and real data sets shows the significant benefits provided by our scheme as compared to existing methods.
Hina A. Khan, Mohamed A. Sharaf, Abdullah M. Albarrak
SSDBM2
2014 Predicting Elections from Social Networks Based on Sub-event Detection and Sentiment Analysis
Sayan Unankard, Xue Li 0001, Mohamed A. Sharaf
WISE (2)3
2014 Quality-aware schedulers for weak consistency key-value data stores
Chen Xu 0001, Mohamed A. Sharaf, Xiaofang Zhou 0001, Aoying Zhou
Distributed Parallel Databases2
2014 A framework for data quality aware query systems
Naiem Khodabandehloo Yeganeh, Shazia Sadiq, Mohamed A. Sharaf
Inf. Syst.3
2014 CoRE: A Context-Aware RelationExtraction Method for Relation Completion
abstract
We identify Relation Completion (RC) as one recurring problem that is central to the success of novel big data applications such as Entity Reconstruction and Data Enrichment.Given a semantic relation R, RC attempts at linking entity pairs between two entity lists under the relation R. To accomplish the RC goals, we propose to formulate search queries for each query entity α based on some auxiliary information, so that to detect its target entity β from the set of retrieved documents.For instance, a Pattern-based method (PaRE) uses extracted patterns as the auxiliary information in formulating search queries.However, high-quality patterns may decrease the probability of finding suitable target entities.As an alternative, we propose CoRE method that uses context terms learned surrounding the expression of a relation as the auxiliary information in formulating queries.The experimental results based on several real-world web data collections demonstrate that CoRE reaches a much higher accuracy than PaRE for the purpose of RC.
Zhixu Li, Mohamed A. Sharaf, Laurianne Sitbon, Xiaoyong Du 0001, Xiaofang Zhou 0001
IEEE Trans. Knowl. Data Eng.2
2014 A web-based approach to data imputation
Zhixu Li, Mohamed A. Sharaf, Laurianne Sitbon, Shazia Sadiq, Marta Indulska, Xiaofang Zhou 0001
World Wide Web2
2013 Location-Based Emerging Event Detection in Social Networks
Sayan Unankard, Xue Li 0001, Mohamed A. Sharaf
APWeb3
2013 Scalable diversification of multiple search results
abstract
The explosion of big data emphasizes the need for scalable data diversification, especially for applications based on web, scientific, and business databases. However, achieving effective diversification in a multi-user environment is a rather challenging task due to the inherent high processing costs of current data diversification techniques. In this paper, we address the concurrent diversification of multiple search results using various approximation techniques that provide orders of magnitude reductions in processing cost, while maintaining comparable quality of diversification as compared to sequential methods. Our extensive experimental evaluation shows the scalability exhibited by our proposed methods under various workload settings.
Hina A. Khan, Marina Drosou, Mohamed A. Sharaf
CIKM3
2013 Adaptive Query Scheduling in Key-Value Data Stores
Chen Xu 0001, Mohamed A. Sharaf, Minqi Zhou, Aoying Zhou, Xiaofang Zhou 0001
DASFAA (1)2
2013 DoS: an efficient scheme for the diversification of multiple search results
abstract
Data diversification provides users with a concise and meaningful view of the results returned by search queries. In addition to taming the information overload, data diversification also provides the benefits of reducing data communication costs as well as enabling data exploration. The explosion of big data emphasizes the need for data diversification in modern data management platforms, especially for applications based on web, scientific, and business databases. Achieving effective diversification, however, is rather a challenging task due to the inherent high processing costs of current data diversification techniques. This challenge is further accentuated in a multi-user environment, in which multiple search queries are to be executed and diversified concurrently. In this paper, we propose the DoS scheme, which addresses the problem of scalable diversification of multiple search results. Our experimental evaluation shows the scalability exhibited by DoS under various workload settings, and the significant benefits it provides compared to sequential methods.
Hina A. Khan, Marina Drosou, Mohamed A. Sharaf
SSDBM3
2012 Efficient buffer management for piecewise linear representation of multiple data streams
abstract
Piecewise Linear Representation (PLR) has been a widely used method for approximating data streams in the form of compact line segments. The buffer-based approach to PLR enables a semi-global approximation which relies on the aggregated processing of batches of streamed data so that to adjust and improve the approximation results. However, one challenge towards applying the buffer-based approach is allocating the necessary memory resources for stream buffering. This challenge is further complicated in a multi-stream environment where multiple data streams are competing for the available memory resources, especially in resource-constrained systems such as sensors and mobile devices.
Qing Xie 0002, Jia Zhu 0003, Mohamed A. Sharaf, Xiaofang Zhou 0001, Chaoyi Pang
CIKM3
2012 Three-Level Processing of Multiple Aggregate Continuous Queries
abstract
Aggregate Continuous Queries (ACQs) are both a very popular class of Continuous Queries (CQs) and also have a potentially high execution cost. As such, optimizing the processing of ACQs is imperative for Data Stream Management Systems (DSMSs) to reach their full potential in supporting (critical) monitoring applications. For multiple ACQs that vary in window specifications and pre-aggregation filters, existing multiple ACQs optimization schemes assume a processing model where each ACQ is computed as a final-aggregation of a sub-aggregation. In this paper, we propose a novel processing model for ACQs, called Tri Ops, with the goal of minimizing the repetition of operator execution at the sub-aggregation level. We also propose Tri Weave, a Tri Ops-aware multi-query optimizer. We analytically and experimentally demonstrate the performance gains of our proposed schemes which shows their superiority over alternative schemes. Finally, we generalize Tri Weave to incorporate the classical subsumption-based multi-query optimization techniques.
Shenoda Guirguis, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis
ICDE2
2012 RFID Mutual Authentication and Secret Update Protocol for Low-Cost Tags
abstract
Low-cost RFID will dominate the industry as a replacement of the barcode tags. RFID tags had been designed for the purpose of automatic identification and tracking. Therefore, RFID tags could violate their owners' privacy and security. Hence, it becomes a necessity to come up with an RFID protocol that meets security (e.g. mutual authentication) and privacy goals. Recently, the Journal of Computer Communications published a paper by Song and Mitchell (2011) where the authors proposed a mutual authentication protocol for RFID system. This protocol has fundamental shortcomings that can be taken advantage by an adept adversary. We shed light on these flaws and later on we modified SM's scheme to fix these vulnerabilities.
Mohamed A. Sharaf
TrustCom1
2012 WebPut: Efficient Web-Based Data Imputation
Zhixu Li, Mohamed A. Sharaf, Laurianne Sitbon, Shazia Sadiq, Marta Indulska, Xiaofang Zhou 0001
WISE2
2012 On the Prediction of Re-tweeting Activities in Social Networks - A Report on WISE 2012 Challenge
Sayan Unankard, Ling Chen 0004, Sen Wang 0001, Zi Huang, Mohamed A. Sharaf, Xue Li 0001
WISE6
2011 Optimized processing of multiple aggregate continuous queries
abstract
Data Streams Management Systems are designed to support monitoring applications, which require the processing of hundreds of Aggregate Continuous Queries (ACQs). These ACQs typically have different time granularities, with possibly different selection predicates and group-by attributes. In order to achieve scalability in the presence of heavy workloads, in this paper, we introduce the concept of 'Weaveability' as an indicator of the potential gains of sharing the processing of ACQs. We then propose Weave Share, a cost-based optimizer that exploits weaveability to optimize the shared processing of ACQs. Our experimental analysis shows that Weave Share outperforms the alternative sharing schemes generating up to four orders of magnitude better quality plans. Finally, we describe a practical implementation of the Weave Share optimizer.
Shenoda Guirguis, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis
CIKM2
2011 Optimizing the Energy Consumption of Continuous Query Processing with Mobile Clients
abstract
Complex event detection over data streams has become ubiquitous through the widespread use of sensors, wireless connectivity and the wide variety of end-user mobile devices. Typically, such event detection is carried out by a data stream management system executing continuous queries (CQs), registered by the users. In this paper, we consider the situation where the results of the CQs, which are in the form of individual data streams, are disseminated to the users' hand-held, battery-operated devices over a shared broadcast medium. In order to reduce the overall energy consumption of the mobile devices, we propose Bose*, a power-aware query operator placement algorithm that determines which part of a CQ plan should be executed at the data stream management system and which part should be executed at the mobile device. Bose*'s effectiveness in reducing energy consumption, as well as response time under specific conditions, is evaluated using simulation, driven by parameters measured on real mobile devices.
Panayiotis Neophytou, Jesse Szwedko, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis
Mobile Data Management (1)3
2011 Visualization of Energy Consumption of Continuous Query Processing with Mobile Clients
abstract
Complex event detection over data streams has become ubiquitous through the widespread use of sensors, wireless connectivity and the wide variety of end-user mobile devices. Typically, event detection is carried out by a central server executing continuous queries. In this demonstration, we focus on the case where users with mobile devices submit continuous queries (for event detection) to a data stream management server which disseminates the results to the users over a shared broadcast medium. In order to minimize the overall energy consumption of the mobile devices (clients), we have proposed operator placement algorithms that split the processing of each continuous query between the centralized server and the requesting mobile clients, thus trading off energy consumption for communication energy consumption for computation. Specifically, in this demonstration, we present an interactive graphical interface to the inner workings of our three proposed operator placement algorithms, whereby attendees are able to investigate various query plans and the decisions that the algorithms make, as well as visualize the results of these algorithms in terms of client power consumption and response time. Besides being able to step through an algorithm's execution as it considers various operator placement decisions, attendees are able to experiment with different scenarios by customizing the parameters of the query workloads (e.g., changing the selectivities and projectivities of the operators) or the client's profile (e.g., power consumed per unit of time of processing) and examine the impact.
Jesse Szwedko, Panayiotis Neophytou, Panos K. Chrysanthis, Alexandros Labrinidis, Mohamed A. Sharaf
Mobile Data Management (1)5
2009 Adaptive Scheduling of Web Transactions
abstract
In highly interactive dynamic Web database systems, user satisfaction determines their success. In such systems, user requested web pages are dynamically created by executing a number of database queries or Web transactions. In this paper, we model the interrelated transactions generating a web page asworkflowsand quantify the user satisfaction by associating dynamic Web pages withsoft-deadlines. Further, we model the importance of transactions in generating a page by associating different weights to transactions. Using this framework, system success is measured in terms of minimizing the deviation from the deadline (i.e., tardiness) and also minimizing the weighted such deviation (i.e., weighted tardiness). In order to efficiently support the materialization of dynamic Web pages, we proposeASETS*, which is a parameter-free adaptive scheduling algorithm that automatically adapts to, not only system load, but also transactions' characteristics (i.e., interdependencies, deadlines and weights).ASETS* prioritizes the execution of transactions with the objective of minimizing weighted tardiness. It is also capable of balancing the tradeoff between optimizing average- and worst-case performance when needed. The performance advantages ofASETS* are experimentally demonstrated.
Shenoda Guirguis, Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Kirk Pruhs
ICDE2
2009 SLA-Aware Adaptive On-demand Data Broadcasting in Wireless Environments
abstract
In mobile and wireless networks, data broadcasting for popular data items enables the efficient utilization of the limited wireless bandwidth. However, efficient data scheduling schemes are needed to fully exploit the benefits of data broadcasting. This motivated the proposal of several broadcast scheduling policies, which have mostly focused on either minimizing response time, or drop rate when requests are associated with hard deadlines. The inherent inaccuracy of hard deadlines in a dynamic mobile environment motivated us to use Service Level Agreements (SLAs) where a user specifies the utility of data as a function of its arrival time. Moreover, SLAs provide the mobile user with an already familiar quality of service specification from wired environments. Hence, in this paper, we propose SAAB which is an SLA-aware adaptive data broadcast scheduling policy for maximizing the system utility under SLA-based performance measures. To achieve this goal, SAAB considers both the characteristics of disseminated data objects as well as the SLAs associated with them. Additionally, SAAB automatically adjusts to the system workload conditions which enables it to constantly outperform existing on-demand broadcast scheduling policies.
Adrian Daniel Popescu, Mohamed A. Sharaf, Cristiana Amza
Mobile Data Management2
2009 Optimizing i/o-intensive transactions in highly interactive applications
abstract
The performance provided by an interactive online database system is typically measured in terms of meeting certain pre-specified Service Level Agreements (SLAs), with expected transaction latency being the most commonly used type of SLA. This form of SLA acts as a soft deadline for each transaction, and user satisfaction can be measured in terms of minimizing tardiness, that is, the deviation from SLA. This objective is further complicated for I/O-intensive transactions, where the storage system becomes the performance bottleneck. Moreover, common I/O scheduling policies employed by the Operating System with a goal of improving I/O throughput or average latency may run counter to optimizing per-transaction performance since the Operating System is typically oblivious to the application high-level SLA specifications. In this paper, we propose a new SLA-aware policy for scheduling I/O requests of database transactions. Our proposed policy synergistically combines novel deadline-aware scheduling policies for database transactions with features of Operating System scheduling policies designed for improving I/O throughput. This enables our proposed policy to dynamically adapt to workload and consistently provide the best performance.
Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Cristiana Amza
SIGMOD Conference1
2008 Scheduling continuous queries in data stream management systems
abstract
Recently, several policies have been proposed for scheduling multiple Continuous Queries (CQs) in a Data Stream Management System (DSMS). The decision on which policy to use plays an important role in shaping the percieved online performance provided by the DSMS. In this tutorial, we provide an overview of different policies employed by current CQ schedulers and the performance goals optimized by these policies. Further, we discuss the salient properties of CQs conisdered by current policies as well as the efficent implementation of such policies into CQ schedulers. Finally, we present future research directions and open problems in CQ scheduling.
Mohamed A. Sharaf, Alexandros Labrinidis, Panos K. Chrysanthis
Proc. VLDB Endow.1
2008 Dynamic partitioning of the cache hierarchy in shared data centers
abstract
Due to the imperative need to reduce the management costs of large data centers, operators multiplex several concurrent database applications on a server farm connected to shared network attached storage. Determining and enforcing per-application resource quotas in the resulting cache hierarchy, on the fly, poses a complex resource allocation problem spanning the database server and the storage server tiers. This problem is further complicated by the need to provide strict Quality of Service (QoS) guarantees to hosted applications. In this paper, we design and implement a novel coordinated partitioning technique of the database buffer pool and storage cache between applications for any given cache replacement policy and per-application access pattern. We use statistical regression to dynamically determine the mapping between cache quota settings and the resulting per-application QoS. A resource controller embedded within the database engine actuates the partitioning of the two-level cache, converging towards the configuration with maximum application utility, expressed as the service provider revenue in that configuration, based on a set of latency sample points. Our experimental evaluation, using the MySQL database engine, a server farm with consolidated storage, and two e-commerce benchmarks, shows the effectiveness of our technique in enforcing application QoS, as well as maximizing the revenue of the service provider in shared server farms.
Gokul Soundararajan, Jin Chen 0006, Mohamed A. Sharaf, Cristiana Amza
Proc. VLDB Endow.3
2008 Algorithms and metrics for processing multiple heterogeneous continuous queries
abstract
The emergence of monitoring applications has precipitated the need for Data Stream Management Systems (DSMSs), which constantly monitor incoming data feeds (through registered continuous queries), in order to detect events of interest. In this article, we examine the problem of how to schedule multiple Continuous Queries (CQs) in a DSMS to optimize different Quality of Service (QoS) metrics. We show that, unlike traditional online systems, scheduling policies in DSMSs that optimize for average response time will be different from policies that optimize for average slowdown, which is a more appropriate metric to use in the presence of a heterogeneous workload. Towards this, we propose policies to optimize for the average-case performance for both metrics. Additionally, we propose a hybrid scheduling policy that strikes a fine balance between performance and fairness, by looking at both the average- and worst-case performance, for both metrics. We also show how our policies can be adaptive enough to handle the inherent dynamic nature of monitoring applications. Furthermore, we discuss how our policies can be efficiently implemented and extended to exploit sharing in optimized multi-query plans and multi-stream CQs. Finally, we experimentally show using real data that our policies consistently outperform currently used ones.
Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Kirk Pruhs
ACM Trans. Database Syst.1
2006 Efficient Scheduling of Heterogeneous Continuous Queries
Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis, Kirk Pruhs
VLDB1
2005 Preemptive rate-based operator scheduling in a data stream management system
abstract
Summary form only given. Data stream management systems are being developed to process continuous queries over multiple data streams. These continuous queries are typically used for monitoring purposes where the detection of an event might trigger a sequence of actions or the execution of a set of specified tasks. Such events are identified by tuples produced by a query and hence, it is important to produce the available portions of a query result as early as possible. A core element for improving the interactive performance of a continuous query is the operator scheduler. An operator scheduler is particularly important when the processing requirements and the productivity of different streams are highly skewed. The need for an operator scheduler becomes even more crucial when tuples from different streams arrive asynchronously. To meet these needs, we are proposing a preemptive rate-based scheduling policy that handles the asynchronous nature of tuple arrival and the heterogeneity in the query plan. Experimental results show the significant improvements provided by our proposed policy.
Mohamed A. Sharaf, Panos K. Chrysanthis, Alexandros Labrinidis
AICCSA1
2005 Freshness-Aware Scheduling of Continuous Queries in the Dynamic Web
Mohamed A. Sharaf, Alexandros Labrinidis, Panos K. Chrysanthis, Kirk Pruhs
WebDB1
2004 On-Demand Data Broadcasting for Mobile Decision Making
Mohamed A. Sharaf, Panos K. Chrysanthis
Mob. Networks Appl.1
2004 Balancing energy efficiency and quality of aggregate data in sensor networks
Mohamed A. Sharaf, Jonathan Beaver, Alexandros Labrinidis, Panos K. Chrysanthis
VLDB J.1
2003 An Optimized Multicast-based Data Dissemination Middleware
abstract
A major problem on the Internet is the scalable dissemination of information. This problem is particularly acute exactly at the time when the scalability of data delivery is most important. One proposed solution to this scalability problem is to use multicast communication. However, allowing multicast communication introduces many nontrivial data management problems, such as caching, consistency, and scheduling. We have built a middleware that unifies and extends state-of-the-art data management methods and algorithms into one software distribution. Its flexible and extensible architecture is built from individual components that can be selected or replaced depending on the underlying multicast transport mechanism or on the application needs. Particular care has gone into the design of the algorithms to optimize the user-perceived level of service. We demonstrate our middleware within the context of the RODS application.
Wenhui Zhang 0002, Vincenzo Liberatore, Vince Penkrot, Jonathan Beaver, Mohamed A. Sharaf, Siddhartha Roychowdhury, Panos K. Chrysanthis, Kirk Pruhs
ICDE6
2003 Efficient Dissemination of Aggregate Data over the Wireless Web
Mohamed A. Sharaf, Yannis Sismanis, Alexandros Labrinidis, Panos K. Chrysanthis, Nick Roussopoulos
WebDB1
2002 Semantic-based delivery of OLAP summary tables in wireless environments
abstract
With the rapid growth in mobile and wireless technologies and the availability, pervasiveness and cost effectiveness of wireless networks, mobile computers are quickly becoming the normal front-end devices for accessing enterprise data. In this paper, we are addressing the issue of efficient delivery of business decision support data in the form of summary tables to mobile clients equipped with OLAP front-end tools. Towards this, we propose a new on-demand scheduling algorithm, called SBS, that exploits both the derivation semantics among OLAP summary tables and the mobile clients' capabilities of executing simple SQL queries. It maximizes the aggregated data sharing between clients and reduces the broadcast length compared to the already existing techniques. The degree of aggregation can be tuned to control the tradeoff between access time and energy consumption. Further, the proposed scheme adapts well to different request rates, access patterns and data distributions. The algorithm effectiveness with respect to access time and power consumption is evaluated using simulation.
Mohamed A. Sharaf, Panos K. Chrysanthis
CIKM1