Mohamed E. Khalefa

dblp:35/7025 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
1since 2021 · last 2021
0009-0007-3123-3527ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 14 · 5 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Query processing and optimization · 57% Spatial and temporal data management · 16% Data mining · 6%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
preference query
0.652013
Flexible and extensible preference evaluation in database systems · ACM Trans. Database Syst. 2013
PrefJoin: An efficient preference-aware join operator · ICDE 2011
CareDB: A Context and Preference-Aware Location-Based Database System · Proc. VLDB Endow. 2010
Information retrieval › evaluation › user-oriented evaluation
preference-based evaluation
0.212013
Flexible and extensible preference evaluation in database systems · ACM Trans. Database Syst. 2013
Data mining › predictive modeling
forecasting
0.112012
Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012
Spatial and temporal data management
time series data management
0.112012
Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012
Spatial and temporal data management › time series data management
time series query
0.112012
Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012
Query processing and optimization
early pruning
0.112011
PrefJoin: An efficient preference-aware join operator · ICDE 2011
Query processing and optimization › join processing
multi-way join
0.112011
On Producing High and Early Result Throughput in Multijoin Query Plans · IEEE Trans. Knowl. Data Eng. 2011
Query processing and optimization › preference query
skyline query
0.122010
Skyline Query Processing for Incomplete Data · ICDE 2008
FlexPref: A framework for extensible preference evaluation in database systems · ICDE 2010
Spatial and temporal data management › location-based services
location-based query
0.112010
CareDB: A Context and Preference-Aware Location-Based Database System · Proc. VLDB Endow. 2010
Database system architecture and tuning › extensible database system
query engine extensibility
0.112010
FlexPref: A framework for extensible preference evaluation in database systems · ICDE 2010
Data stream processing › continuous query processing
early result production
0.112008
PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008
Query processing and optimization
join processing
0.112008
PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008
Data integration and cleaning
missing data
0.112008
Skyline Query Processing for Incomplete Data · ICDE 2008
Query processing and optimization › join processing › join algorithms
non-blocking join
0.112008
PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008
Recommender systems › user modeling
user preference modeling
0.012013
Flexible and extensible preference evaluation in database systems · ACM Trans. Database Syst. 2013
Indexing and storage engines
hierarchical index
0.012012
Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012
Query processing and optimization
adaptive query processing
0.012008
PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008
Data mining › pattern mining › pruning
search space pruning
0.012008
Skyline Query Processing for Incomplete Data · ICDE 2008

Methods — techniques the papers use, named apart from their topics

flushing policy · 0.2preference operators · 0.2statistical forecasting · 0.1model-based representation · 0.1symmetric hash join · 0.1meta data association · 0.1local pruning · 0.1k-dominance · 0.1function registration · 0.1extensible framework · 0.1
YearPublicationVenuePosition
2021 Data-driven Distributionally Robust Optimization For Vehicle Balancing of Mobility-on-Demand Systems
abstract
With the transformation to smarter cities and the development of technologies, a large amount of data is collected from sensors in real time. Services provided by ride-sharing systems such as taxis, mobility-on-demand autonomous vehicles, and bike sharing systems are popular. This paradigm provides opportunities for improving transportation systems’ performance by allocating ride-sharing vehicles toward predicted demand proactively. However, how to deal with uncertainties in the predicted demand probability distribution for improving the average system performance is still a challenging and unsolved task. Considering this problem, in this work, we develop a data-driven distributionally robust vehicle balancing method to minimize the worst-case expected cost. We design efficient algorithms for constructing uncertainty sets of demand probability distributions for different prediction methods and leverage a quad-tree dynamic region partition method for better capturing the dynamic spatial-temporal properties of the uncertain demand. We then derive an equivalent computationally tractable form for numerically solving the distributionally robust problem. We evaluate the performance of the data-driven vehicle balancing algorithm under different demand prediction and region partition methods based on four years of taxi trip data for New York City (NYC). We show that the average total idle driving distance is reduced by 30% with the distributionally robust vehicle balancing method using quad-tree dynamic region partitions, compared with vehicle balancing methods based on static region partitions without considering demand uncertainties. This is about a 60-million-mile or a 8-million-dollar cost reduction annually in NYC.
Fei Miao, Sihong He, Lynn Pepin, Shuo Han 0002, Abdeltawab M. Hendawi, Mohamed E. Khalefa, John A. Stankovic, George J. Pappas
ACM Trans. Cyber Phys. Syst.6
2016 A vision for micro and macro location aware services
abstract
A few decades ago, the Internet was created. Since then, searching for information and services has increased exponentially. With the introduction of GPS-enabled devices, a special type of search appeared offering location-aware services. These services customize search results based on users' location. This includes, but not limited to, (1) service finding, e.g., "find the nearest pizza restaurant", (2) routing, e.g., "obtain the shortest path from a user's home to the airport", (3) transportation, e.g., "what are the bus links to get a user from downtown to the mall", and (4) monitoring, e.g., "alert a parent if their child school-bus deviates from its regular route". Though new hardware and software technologies such as smart watches, voice search, and big-data platforms have been introduced and widely used, each single type of the above services has benefited very little from these technologies. On the local level of each service (the micro level), a full-fledged view is still missing. On the global level of all service types (the macro level), all Location-aware services are still acting as isolated islands and a global optimized service is not available. This paper presents our vision of how to provide an integrated macro location-aware service that acts harmoniously, and how each micro service can be further improved by better incorporation of novel technologies. We also overview the key challenges associated with these suggested improvements. Then, we highlight the potential value-added by the application of our vision.
Abdeltawab M. Hendawi, Mohamed E. Khalefa, Harry Liu, Mohamed H. Ali, John A. Stankovic
SIGSPATIAL/GIS2
2016 Demonstrating KDBMS: A Knowledge-based Database Management System
abstract
We demonstrate a KDBMS, a prototype system which seamlessly integrates Knowledge base and DBMS. While state-of-the-art approaches, i.e., Ontology-based data access, denoted as OBDA, use ontologies to only query data stored in relational databases using SPARQL. In this demo, we present a high level description of the proposed system, introduce a new knowledge-based query language, denoted as KQL, and highlight some query optimization opportunities by employing knowledge across database layers in query optimization, and query processing, while ease the administrating for a complex database schema.
Mohamed E. Khalefa, Sameh S. El-Atawy
SSDBM1
2013 Flexible and extensible preference evaluation in database systems
Justin J. Levandoski, Ahmed Eldawy, Mohamed F. Mokbel, Mohamed E. Khalefa
ACM Trans. Database Syst.4
2012 Aggregating and Disaggregating Flexibility Objects
Laurynas Siksnys, Mohamed E. Khalefa, Torben Bach Pedersen
SSDBM2
2012 Model-based Integration of Past & Future in TimeTravel
abstract
We demonstrate TimeTravel, an efficient DBMS system for seamless integrated querying of past and (forecasted) future values of time series, allowing the user to view past and future values as one joint time series. This functionality is important for advanced application domain like energy. The main idea is to compactly represent time series as models. By using models, the TimeTravel system answers queries approximately on past and future data with error guarantees (absolute error and confidence) one order of magnitude faster than when accessing the time series directly. In addition, it efficiently supports exact historical queries by only accessing relevant portions of the time series. This is unlike existing approaches, which access the entire time series to exactly answer the query. To realize this system, we propose a novel hierarchical model index structure. As real-world time series usually exhibits seasonal behavior, models in this index incorporate seasonality. To construct a hierarchical model index, the user specifies seasonality period, error guarantees levels, and a statistical forecast method. As time proceeds, the system incrementally updates the index and utilizes it to answer approximate and exact queries. TimeTravel is implemented into PostgreSQL, thus achieving complete user transparency at the query level. In the demo, we show the easy building of a hierarchical model index for a real-world time series and the effect of varying the error guarantees on the speed up of approximate and exact queries.
Mohamed E. Khalefa, Ulrike Fischer, Torben Bach Pedersen, Wolfgang Lehner
Proc. VLDB Endow.1
2011 PrefJoin: An efficient preference-aware join operator
abstract
Preference queries are essential to a wide spectrum of applications including multi-criteria decision-making tools and personalized databases. Unfortunately, most of the evaluation techniques for preference queries assume that the set of preferred attributes are stored in only one relation, waiving on a wide set of queries that include preference computations over multiple relations. This paper presents PrefJoin, an efficient preference-aware join query operator, designed specifically to deal with preference queries over multiple relations. PrefJoin consists of four main phases: Local Pruning, Data Preparation, Joining, and Refining that filter out, from each input relation, those tuples that are guaranteed not to be in the final preference set, associate meta data with each non-filtered tuple that will be used to optimize the execution of the next phases, produce a subset of join result that are relevant for the given preference function, and refine these tuples respectively. An interesting characteristic of PrefJoin is that it tightly integrates preference computation with join hence we can early prune those tuples that are guaranteed not to be an answer, and hence it saves significant unnecessary computations cost. PrefJoin supports a variety of preference function including skyline, multi-objective and k-dominance preference queries. We show the correctness of PrefJoin. Experimental evaluation based on a real system implementation inside PostgreSQL shows that PrefJoin consistently achieves from one to three orders of magnitude performance gain over its competitors in various scenarios.
Mohamed E. Khalefa, Mohamed F. Mokbel, Justin J. Levandoski
ICDE1
2011 On Producing High and Early Result Throughput in Multijoin Query Plans
abstract
This paper introduces an efficient framework for producing high and early result throughput in multijoin query plans. While most previous research focuses on optimizing for cases involving a single join operator, this work takes a radical step by addressing query plans with multiple join operators. The proposed framework consists of two main methods, a flush algorithm and operator state manager. The framework assumes a symmetric hash join, a common method for producing early results, when processing incoming data. In this way, our methods can be applied to a group of previous join operators (optimized for single-join queries) when taking part in multijoin query plans. Specifically, our framework can be applied by 1) employing a new flushing policy to write in-memory data to disk, once memory allotment is exhausted, in a way that helps increase the probability of producing early result throughput in multijoin queries, and 2) employing a state manager that adaptively switches operators in the plan between joining in-memory data and disk-resident data in order to positively affect the early result throughput. Extensive experimental results show that the proposed methods outperform the state-of-the-art join operators optimized for both single and multijoin query plans.
Justin J. Levandoski, Mohamed E. Khalefa, Mohamed F. Mokbel
IEEE Trans. Knowl. Data Eng.2
2010 Skyline query processing for uncertain data
abstract
Recently, several research efforts have addressed answering skyline queries efficiently over large datasets. However, this research lacks methods to compute these queries over uncertain data, where uncertain values are represented as a range. In this paper, we define skyline queries over continuous uncertain data, and propose a novel, efficient framework to answer these queries. Query answers are probabilistic, where each object is associated with a probability value of being a query answer. Typically, users specify a probability threshold, that each returned object must exceed, and a tolerance value that defines the allowed error margin in probability calculation to reduce the computational overhead. Our framework employs an efficient two-phase query processing algorithm.
Mohamed E. Khalefa, Mohamed F. Mokbel, Justin J. Levandoski
CIKM1
2010 Preference query evaluation over expensive attributes
abstract
Most database systems allow query processing over attributes that are derived at query runtime (e.g., user-defined functions and remote data calls to web services), making them expensive to compute relative to relational data stored in a heap or index. In addition, core support for efficient preference query processing has become an important objective in database systems. This paper addresses an important problem at the intersection of these two query processing objectives: efficient preference query evaluation involving expensive attributes. We explore an efficient framework for processing skyline and multi-objective queries in a database when the data involves a mix of "cheap" and "expensive" attributes. Our solution involves a three-phase approach that evaluates a correct final preference answer while aiming to minimizing the number of expensive attributes computations. Unlike previous works for distributed preference algorithms that assume sorted access over each attribute, our framework assumes expensive attribute requests are stateless, i.e., know nothing previous requests. Thus, the proposed approach is more in line with realistic system architectures. Our framework is implemented inside the query processor of PostgreSQL, and evaluated over both synthetic and real data sets involving computation of expensive attributes over real web-service data (e.g., Microsoft MapPoint).
Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa
CIKM3
2010 FlexPref: A framework for extensible preference evaluation in database systems
abstract
Personalized database systems give users answers tailored to their personal preferences. While numerous preference evaluation methods for databases have been proposed (e.g., skyline, top-k, k-dominance, k-frequency), the implementation of these methods at the core of a database system is a double-edged sword. Core implementation provides efficient query processing for arbitrary database queries, however this approach is not practical as each existing (and future) preference method requires a custom query processor implementation. To solve this problem, this paper introduces FlexPref, a framework for extensible preference evaluation in database systems. FlexPref, implemented in the query processor, aims to support a wide-array of preference evaluation methods in a single extensible code base. Integration with FlexPref is simple, involving the registration of only three functions that capture the essence of the preference method. Once integrated, the preference method ¿lives¿ at the core of the database, enabling the efficient execution of preference queries involving common database operations. To demonstrate the extensibility of FlexPref, we provide case studies showing the implementation of three database operations (single table access, join, and sorted list access) and five state-of-the-art preference evaluation methods (top-k, skyline, k-dominance, top-k dominance, and k-frequency). We also experimentally study the strengths and weaknesses of an implementation of FlexPef in PostgreSQL over a range of single-table and multi-table preference queries.
Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa
ICDE3
2010 A demonstration of FlexPref: extensible preference evaluation inside the DBMS engine
abstract
This demonstration presents FlexPref, a framework implemented inside the DBMS query processor that enables efficient and extensible preference query processing. FlexPref provides query processing support inside the database engine for a wide-array of preference evaluation methods (e.g., skyline, top-k, k-dominance, k-frequency) in a single extensible code base. Integration with FlexPref is simple, involving the registration of only three functions that capture the essence of the preference method. Once integrated, the preference method "lives" at the core of the database, enabling the efficient execution of preference queries involving common database operations (e.g, selection, join). Functionality of FlexPref, implemented inside PostgreSQL, is demonstrated through the implementation and use of several state-of-the-art preference methods in a real application scenario.
Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa, Venkateshwar R. Korukanti
SIGMOD Conference3
2010 CareDB: A Context and Preference-Aware Location-Based Database System
abstract
We demonstrate CareDB , a context and preference-aware database system. CareDB provides scalable personalized location-based services to users based on their preferences and current surrounding context. Unlike existing location-based database systems that answer queries based solely on proximity in distance, CareDB considers user preferences and various types of context in determining the answer to location-based queries. To this end, CareDB does not aim to define new location-based queries, instead, it aims to redefine the answer of existing location-based queries. To achieve its goals, CareDB has several distinguishing characteristics that revolve around a generic and extensible preference and context-aware query processing framework that addresses (a) scalable, efficient preference joins, (b) gracefully handling contextual attributes that are expensive to derive, and (c) support for uncertain attributes.
Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa
Proc. VLDB Endow.3
2008 Skyline Query Processing for Incomplete Data
abstract
Recently, there has been much interest in processing skyline queries for various applications that include decision making, personalized services, and search pruning. Skyline queries aim to prune a search space of large numbers of multi dimensional data items to a small set of interesting items by eliminating items that are dominated by others. Existing skyline algorithms assume that all dimensions are available for all data items. This paper goes beyond this restrictive assumption as we address the more practical case of involving incomplete data items (i.e., data items missing values in some of their dimensions). In contrast to the case of complete data where the dominance relation is transitive, incomplete data suffer from non-transitive dominance relation which may lead to a cyclic dominance behavior. We first propose two algorithms, namely, "Replacement" and "Bucket" that use traditional skyline algorithms for incomplete data. Then, we propose the "ISkyline" algorithm that is designed specifically for the case of incomplete data. The "ISkyline" algorithm employs two optimization techniques, namely, virtual points and shadow skylines to tolerate cyclic dominance relations. Experimental evidence shows that the "ISkyline" algorithm significantly outperforms variations of traditional skyline algorithms.
Mohamed E. Khalefa, Mohamed F. Mokbel, Justin J. Levandoski
ICDE1
2008 PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans
abstract
This paper introduces an efficient algorithm for Producing Early Results in Multi-join query plans (PermJoin, for short). While most previous research focuses only on the case of a single join operator, PermJoin takes a radical step by addressing query plans with multiple join operators. PermJoin is optimized to maximize the early overall throughput and to adapt to fluctuations in data arrival rates. PermJoin is a non- blocking operator that is capable of producing join results even if one or more data sources are blocked due to slow or bursty network behavior. Furthermore, PermJoin distinguishes itself from all previous techniques as it: (1) employs a new flushing policy to write in-memory data to disk, once memory allotment is exhausted, in a way that helps increase the probability of producing early result throughput in multi-join queries, and (2) employs a novel state manager module that adaptively switches operators between joining in-memory data and disk-resident data in order to maximize overall throughput.
Justin J. Levandoski, Mohamed E. Khalefa, Mohamed F. Mokbel
ICDE2