Zuhair Khayyat

dblp:128/5698 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
0since 2021 · last 2019
0000-0003-3650-6997ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 7 · 4 first-authorSystems, architecture and hardware · 2 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
6 papers
Query processing and optimization · 52% Distributed and cloud data management · 21% Data integration and cleaning · 16%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Parallel and multicore computing · 79% Distributed systems · 14% Performance modeling and evaluation · 7%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization › join processing
inequality join
0.522017
Fast and scalable inequality joins · VLDB J. 2017
Lightning Fast and Space Efficient Inequality Joins · Proc. VLDB Endow. 2015
Query processing and optimization
join processing
0.522017
Fast and scalable inequality joins · VLDB J. 2017
Lightning Fast and Space Efficient Inequality Joins · Proc. VLDB Endow. 2015
Parallel and multicore computing
load balancing
0.422016
Scalemine: scalable parallel frequent subgraph mining in a single large graph · SC 2016
Mizan: a system for dynamic load balancing in large-scale graph processing · EuroSys 2013
Distributed and cloud data management
distributed query processing
0.422017
A Survey and Experimental Comparison of Distributed SPARQL Engines for Very Large RDF Data · Proc. VLDB Endow. 2017
Lightning Fast and Space Efficient Inequality Joins · Proc. VLDB Endow. 2015
Distributed and cloud data management
distributed RDF processing
0.312017
A Survey and Experimental Comparison of Distributed SPARQL Engines for Very Large RDF Data · Proc. VLDB Endow. 2017
Query processing and optimization › join processing
scalable join processing
0.312017
Fast and scalable inequality joins · VLDB J. 2017
Query processing and optimization › query optimization › distributed query optimization
cross-platform query optimization
0.212016
Rheem: Enabling Multi-Platform Task Execution · SIGMOD Conference 2016
Parallel and multicore computing
graph processing
0.212016
Scalemine: scalable parallel frequent subgraph mining in a single large graph · SC 2016
Parallel and multicore computing › graph processing
parallel graph mining
0.212016
Scalemine: scalable parallel frequent subgraph mining in a single large graph · SC 2016
Data integration and cleaning › data preprocessing
data cleaning
0.212015
BigDansing: A System for Big Data Cleansing · SIGMOD Conference 2015
Data integration and cleaning › data quality
data quality rules
0.212015
BigDansing: A System for Big Data Cleansing · SIGMOD Conference 2015
Distributed systems
graph processing systems
0.212013
Mizan: a system for dynamic load balancing in large-scale graph processing · EuroSys 2013
Performance modeling and evaluation
benchmarking
0.112017
A Survey and Experimental Comparison of Distributed SPARQL Engines for Very Large RDF Data · Proc. VLDB Endow. 2017
Data integration and cleaning
data fusion
0.112016
Rheem: Enabling Multi-Platform Task Execution · SIGMOD Conference 2016
Data mining
pattern mining
0.112016
Scalemine: scalable parallel frequent subgraph mining in a single large graph · SC 2016
Query processing and optimization › query optimization › join ordering
join optimization
0.112015
BigDansing: A System for Big Data Cleansing · SIGMOD Conference 2015
Parallel and multicore computing
graph partitioning
0.012013
Mizan: a system for dynamic load balancing in large-scale graph processing · EuroSys 2013

Methods — techniques the papers use, named apart from their topics

experimental evaluation · 0.6two-phase approximate-then-exact mining · 0.5intra-task parallelism · 0.5sorting · 0.3partitioning · 0.3three-layer data processing abstraction · 0.2permutation arrays · 0.2mapreduce · 0.2bloom filter indices · 0.2bit-arrays · 0.2fine-grained vertex migration · 0.2adaptive load balancing · 0.2
YearPublicationVenuePosition
2019 Pivoted Subgraph Isomorphism: The Optimist, the Pessimist and the Realist
Ehab Abdelhamid, Ibrahim Abdelaziz, Zuhair Khayyat, Panos Kalnis
EDBT3
2017 A Survey and Experimental Comparison of Distributed SPARQL Engines for Very Large RDF Data
abstract
Distributed SPARQL engines promise to support very large RDF datasets by utilizing shared-nothing computer clusters. Some are based on distributed frameworks such as MapReduce; others implement proprietary distributed processing; and some rely on expensive preprocessing for data partitioning. These systems exhibit a variety of trade-offs that are not well-understood, due to the lack of any comprehensive quantitative and qualitative evaluation. In this paper, we present a survey of 22 state-of-the-art systems that cover the entire spectrum of distributed RDF data processing and categorize them by several characteristics. Then, we select 12 representative systems and perform extensive experimental evaluation with respect to preprocessing cost, query performance, scalability and workload adaptability, using a variety of synthetic and real large datasets with up to 4.3 billion triples. Our results provide valuable insights for practitioners to understand the trade-offs for their usage scenarios. Finally, we publish online our evaluation framework, including all datasets and workloads, for researchers to compare their novel systems against the existing ones.
Ibrahim Abdelaziz, Razen Al-Harbi, Zuhair Khayyat, Panos Kalnis
Proc. VLDB Endow.3
2017 Errata for "Lightning Fast and Space Efficient Inequality Joins" (PVLDB 8(13): 2074-2085)
abstract
This is in response to recent feedback from some readers, which requires some clarifications regarding our IEJ oin algorithm published in [1]. The feedback revolves around four points: (1) a typo in our illustrating example of the join process; (2) a naming error for the index used by our algorithm to improve the bit array scan; (3) the sort order used in our algorithms; and (4) a missing explanation on how duplicates are handled by our self join algorithm.
Zuhair Khayyat, William Lucia, Meghna Singh, Mourad Ouzzani, Paolo Papotti, Jorge-Arnulfo Quiané-Ruiz, Nan Tang 0001, Panos Kalnis
Proc. VLDB Endow.1
2017 Fast and scalable inequality joins
Zuhair Khayyat, William Lucia, Meghna Singh, Mourad Ouzzani, Paolo Papotti, Jorge-Arnulfo Quiané-Ruiz, Nan Tang 0001, Panos Kalnis
VLDB J.1
2016 Scalemine: scalable parallel frequent subgraph mining in a single large graph
abstract
Frequent Subgraph Mining is an essential operation for graph analytics and knowledge extraction. Due to its high computational cost, parallel solutions are necessary. Existing approaches either suffer from load imbalance, or high communication and synchronization overheads. In this paper we propose ScaleMine; a novel parallel frequent subgraph mining system for a single large graph. ScaleMine introduces a novel two-phase approach. The first phase is approximate; it quickly identifies subgraphs that are frequent with high probability, while collecting various statistics. The second phase computes the exact solution by employing the results of the approximation to achieve good load balance; prune the search space; generate efficient execution plans; and guide intra-task parallelism. Our experiments show that ScaleMine scales to 8,192 cores on a Cray XC40 (12× more than competitors); supports graphs with one billion edges (10× larger than competitors), and is at least an order of magnitude faster than existing solutions.
Ehab Abdelhamid, Ibrahim Abdelaziz, Panos Kalnis, Zuhair Khayyat, Fuad T. Jamour
SC4
2016 Rheem: Enabling Multi-Platform Task Execution
abstract
Many emerging applications, from domains such as healthcare and oil & gas, require several data processing systems for complex analytics. This demo paper showcases system, a framework that provides multi-platform task execution for such applications. It features a three-layer data processing abstraction and a new query optimization approach for multi-platform settings. We will demonstrate the strengths of system by using real-world scenarios from three different applications, namely, machine learning, data cleaning, and data fusion.
Divyakant Agrawal, Mouhamadou Lamine Ba, Laure Berti-Équille, Sanjay Chawla, Ahmed K. Elmagarmid, Hossam M. Hammady, Yasser Idris, Zoi Kaoudi, Zuhair Khayyat, Sebastian Kruse 0001, Mourad Ouzzani, Paolo Papotti, Jorge-Arnulfo Quiané-Ruiz, Nan Tang 0001, Mohammed J. Zaki
SIGMOD Conference9
2015 BigDansing: A System for Big Data Cleansing
abstract
Data cleansing approaches have usually focused on detecting and fixing errors with little attention to scaling to big datasets. This presents a serious impediment since data cleansing often involves costly computations such as enumerating pairs of tuples, handling inequality joins, and dealing with user-defined functions. In this paper, we present BigDansing, a Big Data Cleansing system to tackle efficiency, scalability, and ease-of-use issues in data cleansing. The system can run on top of most common general purpose data processing platforms, ranging from DBMSs to MapReduce-like frameworks. A user-friendly programming interface allows users to express data quality rules both declaratively and procedurally, with no requirement of being aware of the underlying distributed platform. BigDansing takes these rules into a series of transformations that enable distributed computations and several optimizations, such as shared scans and specialized joins operators. Experimental results on both synthetic and real datasets show that BigDansing outperforms existing baseline systems up to more than two orders of magnitude without sacrificing the quality provided by the repair algorithms.
Zuhair Khayyat, Ihab F. Ilyas, Alekh Jindal, Samuel Madden 0001, Mourad Ouzzani, Paolo Papotti, Jorge-Arnulfo Quiané-Ruiz, Nan Tang 0001, Si Yin
SIGMOD Conference1
2015 Lightning Fast and Space Efficient Inequality Joins
abstract
Inequality joins, which join relational tables on inequality conditions, are used in various applications. While there have been a wide range of optimization methods for joins in database systems, from algorithms such as sort-merge join and band join, to various indices such as B + -tree, R * -tree and Bitmap, inequality joins have received little attention and queries containing such joins are usually very slow. In this paper, we introduce fast inequality join algorithms. We put columns to be joined in sorted arrays and we use permutation arrays to encode positions of tuples in one sorted array w.r.t. the other sorted array. In contrast to sort-merge join, we use space efficient bit-arrays that enable optimizations, such as Bloom filter indices, for fast computation of the join results. We have implemented a centralized version of these algorithms on top of PostgreSQL, and a distributed version on top of Spark SQL. We have compared against well known optimization techniques for inequality joins and show that our solution is more scalable and several orders of magnitude faster.
Zuhair Khayyat, William Lucia, Meghna Singh, Mourad Ouzzani, Paolo Papotti, Jorge-Arnulfo Quiané-Ruiz, Nan Tang 0001, Panos Kalnis
Proc. VLDB Endow.1
2013 Mizan: a system for dynamic load balancing in large-scale graph processing
abstract
Pregel [23] was recently introduced as a scalable graph mining system that can provide significant performance improvements over traditional MapReduce implementations. Existing implementations focus primarily on graph partitioning as a preprocessing step to balance computation across compute nodes. In this paper, we examine the runtime characteristics of a Pregel system. We show that graph partitioning alone is insufficient for minimizing end-to-end computation. Especially where data is very large or the runtime behavior of the algorithm is unknown, an adaptive approach is needed. To this end, we introduce Mizan, a Pregel system that achieves efficient load balancing to better adapt to changes in computing needs. Unlike known implementations of Pregel, Mizan does not assume any a priori knowledge of the structure of the graph or behavior of the algorithm. Instead, it monitors the runtime characteristics of the system. Mizan then performs efficient fine-grained vertex migration to balance computation and communication. We have fully implemented Mizan; using extensive evaluation we show that---especially for highly-dynamic workloads---Mizan provides up to 84% improvement over techniques leveraging static graph pre-partitioning.
Zuhair Khayyat, Karim Awara, Amani AlOnazi, Hani Jamjoom, Dan Williams 0001, Panos Kalnis
EuroSys1