VLDB 2026 Research / reviewers in the wild / expert
Mohamed E. Khalefa
dblp:35/7025
· DBLP profile ↗
15ranked-venue papers
5as first author
1since 2021 · last 2021
0009-0007-3123-3527ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 14 · 5 first-authorArtificial intelligence and machine learning · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Query processing and optimization · 57% Spatial and temporal data management · 16% Data mining · 6% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
preference query |
0.6 | 5 | 2013 | Flexible and extensible preference evaluation in database systems · ACM Trans. Database Syst. 2013 PrefJoin: An efficient preference-aware join operator · ICDE 2011 CareDB: A Context and Preference-Aware Location-Based Database System · Proc. VLDB Endow. 2010 |
Information retrieval › evaluation › user-oriented evaluation
preference-based evaluation |
0.2 | 1 | 2013 | Flexible and extensible preference evaluation in database systems · ACM Trans. Database Syst. 2013 |
Data mining › predictive modeling
forecasting |
0.1 | 1 | 2012 | Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012 |
Spatial and temporal data management
time series data management |
0.1 | 1 | 2012 | Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012 |
Spatial and temporal data management › time series data management
time series query |
0.1 | 1 | 2012 | Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012 |
Query processing and optimization
early pruning |
0.1 | 1 | 2011 | PrefJoin: An efficient preference-aware join operator · ICDE 2011 |
Query processing and optimization › join processing
multi-way join |
0.1 | 1 | 2011 | On Producing High and Early Result Throughput in Multijoin Query Plans · IEEE Trans. Knowl. Data Eng. 2011 |
Query processing and optimization › preference query
skyline query |
0.1 | 2 | 2010 | Skyline Query Processing for Incomplete Data · ICDE 2008 FlexPref: A framework for extensible preference evaluation in database systems · ICDE 2010 |
Spatial and temporal data management › location-based services
location-based query |
0.1 | 1 | 2010 | CareDB: A Context and Preference-Aware Location-Based Database System · Proc. VLDB Endow. 2010 |
Database system architecture and tuning › extensible database system
query engine extensibility |
0.1 | 1 | 2010 | FlexPref: A framework for extensible preference evaluation in database systems · ICDE 2010 |
Data stream processing › continuous query processing
early result production |
0.1 | 1 | 2008 | PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008 |
Query processing and optimization
join processing |
0.1 | 1 | 2008 | PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008 |
Data integration and cleaning
missing data |
0.1 | 1 | 2008 | Skyline Query Processing for Incomplete Data · ICDE 2008 |
Query processing and optimization › join processing › join algorithms
non-blocking join |
0.1 | 1 | 2008 | PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008 |
Recommender systems › user modeling
user preference modeling |
0.0 | 1 | 2013 | Flexible and extensible preference evaluation in database systems · ACM Trans. Database Syst. 2013 |
Indexing and storage engines
hierarchical index |
0.0 | 1 | 2012 | Model-based Integration of Past & Future in TimeTravel · Proc. VLDB Endow. 2012 |
Query processing and optimization
adaptive query processing |
0.0 | 1 | 2008 | PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query Plans · ICDE 2008 |
Data mining › pattern mining › pruning
search space pruning |
0.0 | 1 | 2008 | Skyline Query Processing for Incomplete Data · ICDE 2008 |
Methods — techniques the papers use, named apart from their topics
flushing policy · 0.2preference operators · 0.2statistical forecasting · 0.1model-based representation · 0.1symmetric hash join · 0.1meta data association · 0.1local pruning · 0.1k-dominance · 0.1function registration · 0.1extensible framework · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | Data-driven Distributionally Robust Optimization For Vehicle Balancing of Mobility-on-Demand SystemsabstractWith the transformation to smarter cities and the development of technologies, a large amount of data is collected from sensors in real time. Services provided by ride-sharing systems such as taxis, mobility-on-demand autonomous vehicles, and bike sharing systems are popular. This paradigm provides opportunities for improving transportation systems’ performance by allocating ride-sharing vehicles toward predicted demand proactively. However, how to deal with uncertainties in the predicted demand probability distribution for improving the average system performance is still a challenging and unsolved task. Considering this problem, in this work, we develop a data-driven distributionally robust vehicle balancing method to minimize the worst-case expected cost. We design efficient algorithms for constructing uncertainty sets of demand probability distributions for different prediction methods and leverage a quad-tree dynamic region partition method for better capturing the dynamic spatial-temporal properties of the uncertain demand. We then derive an equivalent computationally tractable form for numerically solving the distributionally robust problem. We evaluate the performance of the data-driven vehicle balancing algorithm under different demand prediction and region partition methods based on four years of taxi trip data for New York City (NYC). We show that the average total idle driving distance is reduced by 30% with the distributionally robust vehicle balancing method using quad-tree dynamic region partitions, compared with vehicle balancing methods based on static region partitions without considering demand uncertainties. This is about a 60-million-mile or a 8-million-dollar cost reduction annually in NYC. Fei Miao, Sihong He, Lynn Pepin, Shuo Han 0002, Abdeltawab M. Hendawi, Mohamed E. Khalefa, John A. Stankovic, George J. Pappas |
ACM Trans. Cyber Phys. Syst. | 6 |
| 2016 | A vision for micro and macro location aware servicesabstractA few decades ago, the Internet was created. Since then, searching for information and services has increased exponentially. With the introduction of GPS-enabled devices, a special type of search appeared offering location-aware services. These services customize search results based on users' location. This includes, but not limited to, (1) service finding, e.g., "find the nearest pizza restaurant", (2) routing, e.g., "obtain the shortest path from a user's home to the airport", (3) transportation, e.g., "what are the bus links to get a user from downtown to the mall", and (4) monitoring, e.g., "alert a parent if their child school-bus deviates from its regular route". Though new hardware and software technologies such as smart watches, voice search, and big-data platforms have been introduced and widely used, each single type of the above services has benefited very little from these technologies. On the local level of each service (the micro level), a full-fledged view is still missing. On the global level of all service types (the macro level), all Location-aware services are still acting as isolated islands and a global optimized service is not available. This paper presents our vision of how to provide an integrated macro location-aware service that acts harmoniously, and how each micro service can be further improved by better incorporation of novel technologies. We also overview the key challenges associated with these suggested improvements. Then, we highlight the potential value-added by the application of our vision. Abdeltawab M. Hendawi, Mohamed E. Khalefa, Harry Liu, Mohamed H. Ali, John A. Stankovic |
SIGSPATIAL/GIS | 2 |
| 2016 | Demonstrating KDBMS: A Knowledge-based Database Management SystemabstractWe demonstrate a KDBMS, a prototype system which seamlessly integrates Knowledge base and DBMS. While state-of-the-art approaches, i.e., Ontology-based data access, denoted as OBDA, use ontologies to only query data stored in relational databases using SPARQL. In this demo, we present a high level description of the proposed system, introduce a new knowledge-based query language, denoted as KQL, and highlight some query optimization opportunities by employing knowledge across database layers in query optimization, and query processing, while ease the administrating for a complex database schema. Mohamed E. Khalefa, Sameh S. El-Atawy |
SSDBM | 1 |
| 2013 | Flexible and extensible preference evaluation in database systems
Justin J. Levandoski, Ahmed Eldawy, Mohamed F. Mokbel, Mohamed E. Khalefa |
ACM Trans. Database Syst. | 4 |
| 2012 | Aggregating and Disaggregating Flexibility Objects
Laurynas Siksnys, Mohamed E. Khalefa, Torben Bach Pedersen |
SSDBM | 2 |
| 2012 | Model-based Integration of Past & Future in TimeTravelabstractWe demonstrate TimeTravel, an efficient DBMS system for seamless integrated querying of past and (forecasted) future values of time series, allowing the user to view past and future values as one joint time series. This functionality is important for advanced application domain like energy. The main idea is to compactly represent time series as models. By using models, the TimeTravel system answers queries approximately on past and future data with error guarantees (absolute error and confidence) one order of magnitude faster than when accessing the time series directly. In addition, it efficiently supports exact historical queries by only accessing relevant portions of the time series. This is unlike existing approaches, which access the entire time series to exactly answer the query. To realize this system, we propose a novel hierarchical model index structure. As real-world time series usually exhibits seasonal behavior, models in this index incorporate seasonality. To construct a hierarchical model index, the user specifies seasonality period, error guarantees levels, and a statistical forecast method. As time proceeds, the system incrementally updates the index and utilizes it to answer approximate and exact queries. TimeTravel is implemented into PostgreSQL, thus achieving complete user transparency at the query level. In the demo, we show the easy building of a hierarchical model index for a real-world time series and the effect of varying the error guarantees on the speed up of approximate and exact queries. Mohamed E. Khalefa, Ulrike Fischer, Torben Bach Pedersen, Wolfgang Lehner |
Proc. VLDB Endow. | 1 |
| 2011 | PrefJoin: An efficient preference-aware join operatorabstractPreference queries are essential to a wide spectrum of applications including multi-criteria decision-making tools and personalized databases. Unfortunately, most of the evaluation techniques for preference queries assume that the set of preferred attributes are stored in only one relation, waiving on a wide set of queries that include preference computations over multiple relations. This paper presents PrefJoin, an efficient preference-aware join query operator, designed specifically to deal with preference queries over multiple relations. PrefJoin consists of four main phases: Local Pruning, Data Preparation, Joining, and Refining that filter out, from each input relation, those tuples that are guaranteed not to be in the final preference set, associate meta data with each non-filtered tuple that will be used to optimize the execution of the next phases, produce a subset of join result that are relevant for the given preference function, and refine these tuples respectively. An interesting characteristic of PrefJoin is that it tightly integrates preference computation with join hence we can early prune those tuples that are guaranteed not to be an answer, and hence it saves significant unnecessary computations cost. PrefJoin supports a variety of preference function including skyline, multi-objective and k-dominance preference queries. We show the correctness of PrefJoin. Experimental evaluation based on a real system implementation inside PostgreSQL shows that PrefJoin consistently achieves from one to three orders of magnitude performance gain over its competitors in various scenarios. Mohamed E. Khalefa, Mohamed F. Mokbel, Justin J. Levandoski |
ICDE | 1 |
| 2011 | On Producing High and Early Result Throughput in Multijoin Query PlansabstractThis paper introduces an efficient framework for producing high and early result throughput in multijoin query plans. While most previous research focuses on optimizing for cases involving a single join operator, this work takes a radical step by addressing query plans with multiple join operators. The proposed framework consists of two main methods, a flush algorithm and operator state manager. The framework assumes a symmetric hash join, a common method for producing early results, when processing incoming data. In this way, our methods can be applied to a group of previous join operators (optimized for single-join queries) when taking part in multijoin query plans. Specifically, our framework can be applied by 1) employing a new flushing policy to write in-memory data to disk, once memory allotment is exhausted, in a way that helps increase the probability of producing early result throughput in multijoin queries, and 2) employing a state manager that adaptively switches operators in the plan between joining in-memory data and disk-resident data in order to positively affect the early result throughput. Extensive experimental results show that the proposed methods outperform the state-of-the-art join operators optimized for both single and multijoin query plans. Justin J. Levandoski, Mohamed E. Khalefa, Mohamed F. Mokbel |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2010 | Skyline query processing for uncertain dataabstractRecently, several research efforts have addressed answering skyline queries efficiently over large datasets. However, this research lacks methods to compute these queries over uncertain data, where uncertain values are represented as a range. In this paper, we define skyline queries over continuous uncertain data, and propose a novel, efficient framework to answer these queries. Query answers are probabilistic, where each object is associated with a probability value of being a query answer. Typically, users specify a probability threshold, that each returned object must exceed, and a tolerance value that defines the allowed error margin in probability calculation to reduce the computational overhead. Our framework employs an efficient two-phase query processing algorithm. Mohamed E. Khalefa, Mohamed F. Mokbel, Justin J. Levandoski |
CIKM | 1 |
| 2010 | Preference query evaluation over expensive attributesabstractMost database systems allow query processing over attributes that are derived at query runtime (e.g., user-defined functions and remote data calls to web services), making them expensive to compute relative to relational data stored in a heap or index. In addition, core support for efficient preference query processing has become an important objective in database systems. This paper addresses an important problem at the intersection of these two query processing objectives: efficient preference query evaluation involving expensive attributes. We explore an efficient framework for processing skyline and multi-objective queries in a database when the data involves a mix of "cheap" and "expensive" attributes. Our solution involves a three-phase approach that evaluates a correct final preference answer while aiming to minimizing the number of expensive attributes computations. Unlike previous works for distributed preference algorithms that assume sorted access over each attribute, our framework assumes expensive attribute requests are stateless, i.e., know nothing previous requests. Thus, the proposed approach is more in line with realistic system architectures. Our framework is implemented inside the query processor of PostgreSQL, and evaluated over both synthetic and real data sets involving computation of expensive attributes over real web-service data (e.g., Microsoft MapPoint). Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa |
CIKM | 3 |
| 2010 | FlexPref: A framework for extensible preference evaluation in database systemsabstractPersonalized database systems give users answers tailored to their personal preferences. While numerous preference evaluation methods for databases have been proposed (e.g., skyline, top-k, k-dominance, k-frequency), the implementation of these methods at the core of a database system is a double-edged sword. Core implementation provides efficient query processing for arbitrary database queries, however this approach is not practical as each existing (and future) preference method requires a custom query processor implementation. To solve this problem, this paper introduces FlexPref, a framework for extensible preference evaluation in database systems. FlexPref, implemented in the query processor, aims to support a wide-array of preference evaluation methods in a single extensible code base. Integration with FlexPref is simple, involving the registration of only three functions that capture the essence of the preference method. Once integrated, the preference method ¿lives¿ at the core of the database, enabling the efficient execution of preference queries involving common database operations. To demonstrate the extensibility of FlexPref, we provide case studies showing the implementation of three database operations (single table access, join, and sorted list access) and five state-of-the-art preference evaluation methods (top-k, skyline, k-dominance, top-k dominance, and k-frequency). We also experimentally study the strengths and weaknesses of an implementation of FlexPef in PostgreSQL over a range of single-table and multi-table preference queries. Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa |
ICDE | 3 |
| 2010 | A demonstration of FlexPref: extensible preference evaluation inside the DBMS engineabstractThis demonstration presents FlexPref, a framework implemented inside the DBMS query processor that enables efficient and extensible preference query processing. FlexPref provides query processing support inside the database engine for a wide-array of preference evaluation methods (e.g., skyline, top-k, k-dominance, k-frequency) in a single extensible code base. Integration with FlexPref is simple, involving the registration of only three functions that capture the essence of the preference method. Once integrated, the preference method "lives" at the core of the database, enabling the efficient execution of preference queries involving common database operations (e.g, selection, join). Functionality of FlexPref, implemented inside PostgreSQL, is demonstrated through the implementation and use of several state-of-the-art preference methods in a real application scenario. Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa, Venkateshwar R. Korukanti |
SIGMOD Conference | 3 |
| 2010 | CareDB: A Context and Preference-Aware Location-Based Database SystemabstractWe demonstrate CareDB , a context and preference-aware database system. CareDB provides scalable personalized location-based services to users based on their preferences and current surrounding context. Unlike existing location-based database systems that answer queries based solely on proximity in distance, CareDB considers user preferences and various types of context in determining the answer to location-based queries. To this end, CareDB does not aim to define new location-based queries, instead, it aims to redefine the answer of existing location-based queries. To achieve its goals, CareDB has several distinguishing characteristics that revolve around a generic and extensible preference and context-aware query processing framework that addresses (a) scalable, efficient preference joins, (b) gracefully handling contextual attributes that are expensive to derive, and (c) support for uncertain attributes. Justin J. Levandoski, Mohamed F. Mokbel, Mohamed E. Khalefa |
Proc. VLDB Endow. | 3 |
| 2008 | Skyline Query Processing for Incomplete DataabstractRecently, there has been much interest in processing skyline queries for various applications that include decision making, personalized services, and search pruning. Skyline queries aim to prune a search space of large numbers of multi dimensional data items to a small set of interesting items by eliminating items that are dominated by others. Existing skyline algorithms assume that all dimensions are available for all data items. This paper goes beyond this restrictive assumption as we address the more practical case of involving incomplete data items (i.e., data items missing values in some of their dimensions). In contrast to the case of complete data where the dominance relation is transitive, incomplete data suffer from non-transitive dominance relation which may lead to a cyclic dominance behavior. We first propose two algorithms, namely, "Replacement" and "Bucket" that use traditional skyline algorithms for incomplete data. Then, we propose the "ISkyline" algorithm that is designed specifically for the case of incomplete data. The "ISkyline" algorithm employs two optimization techniques, namely, virtual points and shadow skylines to tolerate cyclic dominance relations. Experimental evidence shows that the "ISkyline" algorithm significantly outperforms variations of traditional skyline algorithms. Mohamed E. Khalefa, Mohamed F. Mokbel, Justin J. Levandoski |
ICDE | 1 |
| 2008 | PermJoin: An Efficient Algorithm for Producing Early Results in Multi-join Query PlansabstractThis paper introduces an efficient algorithm for Producing Early Results in Multi-join query plans (PermJoin, for short). While most previous research focuses only on the case of a single join operator, PermJoin takes a radical step by addressing query plans with multiple join operators. PermJoin is optimized to maximize the early overall throughput and to adapt to fluctuations in data arrival rates. PermJoin is a non- blocking operator that is capable of producing join results even if one or more data sources are blocked due to slow or bursty network behavior. Furthermore, PermJoin distinguishes itself from all previous techniques as it: (1) employs a new flushing policy to write in-memory data to disk, once memory allotment is exhausted, in a way that helps increase the probability of producing early result throughput in multi-join queries, and (2) employs a novel state manager module that adaptively switches operators between joining in-memory data and disk-resident data in order to maximize overall throughput. Justin J. Levandoski, Mohamed E. Khalefa, Mohamed F. Mokbel |
ICDE | 2 |