Mohammed Al-Kateb

dblp:11/820 · DBLP profile ↗
← Back
17ranked-venue papers
11as first author
2since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 10 first-author · 2 since 2021Artificial intelligence and machine learning · 2 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Query processing and optimization · 59% Distributed and cloud data management · 15% Data stream processing · 10%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Performance modeling and evaluation · 70% Cloud and datacenter computing · 21% Storage systems · 9%

Topics — the 17 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query optimization
1.132021
Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage · Proc. VLDB Endow. 2021
Dynamic Statistics Collection in the Teradata Unified Data Architecture · ICDE 2017
Optimizing UNION ALL Join Queries in Teradata · ICDE 2017
Distributed and cloud data management
distributed analytics
0.612022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Data stream processing
streaming analytics
0.612022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Machine learning and data management
in-database machine learning
0.412019
In-database Distributed Machine Learning: Demonstration using Teradata SQL Engine · Proc. VLDB Endow. 2019
Query processing and optimization › query optimization
cost-based optimization
0.312018
Joins over UNION ALL Queries in Teradata®: Demonstration of Optimized Execution · SIGMOD Conference 2018
Query processing and optimization › query optimization › join ordering
join optimization
0.312017
Optimizing UNION ALL Join Queries in Teradata · ICDE 2017
Data models and query languages › semistructured data
semi-structured data model
0.312017
BigBench V2: The New and Improved BigBench · ICDE 2017
Query processing and optimization › query optimization › statistics management
statistics collection
0.312017
Dynamic Statistics Collection in the Teradata Unified Data Architecture · ICDE 2017
Performance modeling and evaluation
benchmarking
0.312017
BigBench V2: The New and Improved BigBench · ICDE 2017
Performance modeling and evaluation › benchmarking › distributed system benchmarking
big data system benchmarking
0.312017
BigBench V2: The New and Improved BigBench · ICDE 2017
Distributed and cloud data management
data partitioning
0.212016
Hybrid Row-Column Partitioning in Teradata · Proc. VLDB Endow. 2016
Query processing and optimization
analytical query processing
0.212022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Cloud and datacenter computing
cloud platform
0.212022
Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels · ICDE 2022
Query processing and optimization › query planning
query plan generation
0.112012
Adaptive optimizations of recursive queries in teradata · SIGMOD Conference 2012
Query processing and optimization › recursive query
recursive query optimization
0.112012
Adaptive optimizations of recursive queries in teradata · SIGMOD Conference 2012
Storage systems › storage engine
storage engine design
0.112016
Hybrid Row-Column Partitioning in Teradata · Proc. VLDB Endow. 2016
Web and social media mining
social network analysis
0.012012
Adaptive optimizations of recursive queries in teradata · SIGMOD Conference 2012

Methods — techniques the papers use, named apart from their topics

distributed computing · 1.1scale factor-based data generation · 0.6projection push · 0.5predicate push · 0.5markup language-based property specification · 0.5neural network training · 0.4linear regression · 0.4query execution feedback · 0.3hash-based join · 0.3cost-based join pushing · 0.3performance study · 0.2
YearPublicationVenuePosition
2022 Analytics at Scale: Evolution at Infrastructure and Algorithmic Levels
abstract
Data Analytics is at the core of almost all modern ap-plications ranging from science and finance to healthcare and web applications. The evolution of data analytics over the last decade has been dramatic - new methods, new tools and new platforms - with no slowdown in sight. This rapid evolution has pushed the boundaries of data analytics along several axis including scalability especially with the rise of distributed infrastructures and the Big Data era, and interoperability with diverse data management systems such as relational databases, Hadoop and Spark. However, many analytic application developers struggle with the challenge of production deployment. Recent experience suggests that it is difficult to deliver modern data analytics with the level of reliability, security and manageability that has been a feature of traditional SQL DBMSs. In this tutorial, we discuss the advances and innovations introduced at both the infrastructure and algorithmic levels, directed at making analytic workloads scale, while paying close attention to the kind of quality of service guarantees different technology provide. We start with an overview of the classical centralized analytical techniques, describing the shift towards distributed analytics over non-SQL infrastructures. We contrast such approaches with systems that integrate analytic functionality inside, above or adjacent to SQL engines. We also explore how Cloud platforms' virtualization capabilities make it easier - and cheaper - for end users to apply these new analytic techniques to their data. Finally, we conclude with the learned lessons and a vision for the near future.
Mohammed Al-Kateb, Mohamed Y. Eltabakh, Awny Al-Omari, Paul G. Brown
ICDE1
2021 Not Black-Box Anymore! Enabling Analytics-Aware Optimizations in Teradata Vantage
abstract
Teradata Vantage is a platform for integrating a broad range of analytical functions and capabilities with the Teradata's SQL engine. One of the main challenges in optimizing the execution of these analytical functions is that many of them are not only black boxes, but also have polymorphic nature, i.e., their behavior and properties may change depending on the invocation context. In this paper, we first demonstrate the inherent complexity in optimizing polymorphic functions, and then present the Vantage's Collaborative Optimizer , which is a cross-platform optimizer designed for optimizing the analytical functions invoked from within the SQL engine. The Collaborative Optimizer is the industry-first effort towards enabling analytics-aware optimizations over polymorphic analytical functions. We present a novel markup language-based approach for expressing the functions' polymorphic properties via a set of well-defined instructions. The Collaborative Optimizer uses these instructions at query time to infer the corresponding properties, and then decide on the applicable optimizations. From several possible optimizations, we showcase two core optimizations, namely "projection push" and "predicate push" , which aim at optimizing the data movement to and from the analytical functions. The experiments using the Teradata-MLE analytical system demonstrate the expressiveness power and flexibility of the proposed markup language. Moreover, benchmark and real-world customer queries show the significant performance gain that the Collaborative Optimizer brings to the Vantage system.
Mohamed Y. Eltabakh, Anantha Subramanian, Awny Al-Omari, Mohammed Al-Kateb, Sanjay Nair, Mahbub Hasan, Wellington Cabrera, Amit Kishore, Snigdha Prasad
Proc. VLDB Endow.4
2020 Cost Estimation Across Heterogeneous SQL-Based Big Data Infrastructures in Teradata IntelliSphere
Kassem Awada, Mohamed Y. Eltabakh, Conrad Tang, Mohammed Al-Kateb, Sanjay Nair, Grace Au
EDBT4
2019 In-database Distributed Machine Learning: Demonstration using Teradata SQL Engine
abstract
Machine learning has enabled many interesting applications and is extensively being used in big data systems. The popular approach - training machine learning models in frameworks like Tensorflow, Pytorch and Keras - requires movement of data from database engines to analytical engines, which adds an excessive overhead on data scientists and becomes a performance bottleneck for model training. In this demonstration, we give a practical exhibition of a solution for the enablement of distributed machine learning natively inside database engines. During the demo, the audience will interactively use Python APIs in Jupyter Notebooks to train multiple linear regression models on synthetic regression datasets and neural network models on vision and sensory datasets directly inside Teradata SQL Engine.
Sandeep Singh Sandha, Wellington Cabrera, Mohammed Al-Kateb, Sanjay Nair, Mani Srivastava 0001
Proc. VLDB Endow.3
2018 Joins over UNION ALL Queries in Teradata®: Demonstration of Optimized Execution
abstract
The UNION ALL set operator is useful for combining data from multiple sources. With the emergence and prevalence of big data ecosystems in which data is typically stored on multiple systems, UNION ALL has become even more important in many analytical queries. In this project, we demonstrate novel cost-based optimization techniques implemented in Teradata Database for join queries involving UNION ALL views and derived tables. Instead of the naive and traditional way of spooling each UNION ALL branch to a common spool prior to performing join operations, which can be prohibitively expensive, we demonstrate new techniques developed in Teradata Database including: 1) Cost-based pushing of joins into UNION ALL branches, 2) Branch grouping strategy prior to join pushing, 3) Geography adjustment of the pushed relations to avoid unnecessary redistribution or duplication, 4) Iterative join decomposition of a pushed join to multiple joins, and 5) Combining multiple join steps into a single multisource join step. In the demonstration, we use the Teradata Visual Explain tool, which offers a rich set of visual rendering capabilities of query plans, the display of various metadata information for each plan step, and several interactive UGI options for end-users.
Mohammed Al-Kateb, Paul Sinclair, Grace Au, Sanjay Nair, Mark Sirek, Mohamed Y. Eltabakh
SIGMOD Conference1
2017 Optimizing UNION ALL Join Queries in Teradata
abstract
The UNION ALL set operator is useful for combining data from multiple sources. With the emergence of big data ecosystems in which data is typically stored on multiple systems, UNION ALL has become even more important. In this paper, we present optimization techniques implemented in Teradata Database for join queries with UNION ALL. Instead of spooling all branches of UNION ALL before performing join operations, we propose cost-based pushing of joins into branches. Join pushing not only addresses the prohibitive cost of spooling all branches, but also helps in exposing more efficient join methods (e.g., direct hash-based joins) which, otherwise, would not be considered by the query optimizer. The geography of relations being pushed to UNION ALL branches is also adjusted to avoid unnecessary redistributions and duplications of data. We conclude the paper with a performance study that demonstrates the impact of the proposed optimization techniques on query performance.
Mohammed Al-Kateb, Paul Sinclair, Alain Crolotte, Grace Au, Sanjay Nair
ICDE1
2017 BigBench V2: The New and Improved BigBench
abstract
Benchmarking Big Data solutions has been gaining a lot of attention from research and industry. BigBench is one of the most popular benchmarks in this area which was adopted by the TPC as TPCx-BB. BigBench, however, has key shortcomings. The structured component of the data model is the same as the TPC-DS data model which is a complex snowflake-like schema. This is contrary to the simple star schema Big Data models in real life. BigBench also treats the semi-structured web-logs more or less as a structured table. In real life, web-logs are modeled as key-value pairs with unknown schema. Specific keys are captured at query time - a process referred to as late binding. In addition, eleven (out of thirty) of the BigBench queries are TPC-DS queries. These queries are complex SQL applied on the structured part of the data model which again is not typical of Big Data workloads. In this paper1, we present BigBench V2 to address the aforementioned limitations of the original BigBench. BigBench V2 is completely independent of TPC-DS with a new data model and an overhauled workload. The new data model has a simple structured data model. Web-logs are modeled as key-value pairs with a substantial and variable number of keys. BigBench V2 mandates late binding by requiring query processing to be done directly on key-value web-logs rather than a pre-parsed form of it. A new scale factor-based data generator is implemented to produce structured tables, key-value semistructured web-logs, and unstructured data. We implemented and executed BigBench V2 on Hive. Our proof of concept shows the feasibility of BigBench V2 and outlines different ways of implementing late binding.
Ahmad Ghazal, Todor Ivanov, Pekka Kostamaa, Alain Crolotte, Ryan Voong, Mohammed Al-Kateb, Waleed Ghazal, Roberto V. Zicari
ICDE6
2017 Dynamic Statistics Collection in the Teradata Unified Data Architecture
abstract
The Unified Data Architecture (UDA) of Teradata is an inclusive multisystem data analytics solution. A key challenge for query optimization under the UDA is to find optimal plans for queries that access data on heterogeneous remote data stores. The challenge comes primarily from the lack of statistics for data stored on remote systems. In this paper, we present techniques implemented in Teradata Database for dynamically collecting statistics on data fetched from a remote system and feeding these statistics back to the query optimizer during query execution. We demonstrate the performance impact of dynamic statistics collection and feedback with experiments conducted on a system that consists of Teradata Database and a remote Hadoop server.
Mohammed Al-Kateb, Paul Sinclair, Alain Crolotte, Linda Rose
ICDE2
2016 Hybrid Row-Column Partitioning in Teradata
abstract
Data partitioning is an indispensable ingredient of database systems due to the performance improvement it can bring to any given mixed workload. Data can be partitioned horizontally or vertically. While some commercial proprietary and open source database systems have one flavor or mixed flavors of these partitioning forms, Teradata Database offers a unique hybrid row-column store solution that seamlessly combines both of these partitioning schemes. The key feature of this hybrid solution is that either row, column, or combined partitions are all stored and handled in the same way internally by the underlying file system storage layer. In this paper, we present the main characteristics and explain the implementation approach of Teradata's row-column store. We also discuss query optimization techniques applicable specifically to partitioned tables. Furthermore, we present a performance study that demonstrates how different partitioning options impact the performance of various queries.
Mohammed Al-Kateb, Paul Sinclair, Grace Au, Carrie Ballinger
Proc. VLDB Endow.1
2014 Adaptive stratified reservoir sampling over heterogeneous data streams
Mohammed Al-Kateb, Byung Suk Lee 0001
Inf. Syst.1
2013 Temporal query processing in Teradata
abstract
The importance of temporal data management is evident by the temporal features recently released in major commercial database systems. In Teradata, the temporal feature is based on the TSQL2 specification. In this paper, we present Teradata's implementation approach for temporal query processing. There are two common approaches to support temporal query processing in a database engine. One is through functional query rewrites to convert a temporal query to a semantically-equivalent non-temporal counterpart, mostly by adding time-based constraints. The other is a native support that implements temporal database operations such as scans and joins directly in the DBMS internals. These approaches have competing pros and cons. The rewrite approach is generally simpler to implement. But it adds a structural complexity to original query, which can pose a potential challenge to query optimizer and cause it to generate sub-optimal plans. A native support is expected to perform better. But it usually involves a higher cost of implementation, maintenance, and extension. We discuss why and describe how Teradata adopted the rewrite approach. In addition, we present an evaluation of our approach through a performance study conducted on a variation of the TPC-H benchmark with temporal tables and queries.
Mohammed Al-Kateb, Ahmad Ghazal, Alain Crolotte, Ramesh Bhashyam, Jaiprakash Chimanchode, Sai Pavan Pakala
EDBT1
2012 An Efficient SQL Rewrite Approach for Temporal Coalescing in the Teradata RDBMS
Mohammed Al-Kateb, Ahmad Ghazal, Alain Crolotte
DEXA (2)1
2012 Adaptive optimizations of recursive queries in teradata
abstract
Recursive queries were introduced as part of ANSI SQL 99 to support processing of hierarchical data typical of air flight schedules, bill-of-materials, data cube dimension hierarchies, and ancestor-descendant information (e.g. XML data stored in relations). Recently, recursive queries have also found extensive use in web data analysis such as social network and click stream data. Teradata implemented recursive queries in V2R6 using static plans whereby a query is executed in multiple iterations, each iteration corresponding to one level of the recursion. Such a static planning strategy may not be optimal since the demographics of intermediate results from recursive iterations often vary to a great extent. Gathering feedback at each iteration could address this problem by providing size estimates to the optimizer which, in turn, can produce an execution plan for the next iteration. However, such a full feedback scheme suffers from lack of pipelining and the inability to exploit global optimizations across the different recursion iterations. In this paper, we propose adaptive optimization techniques that avoid the issues with static as well as full feedback optimization approaches. Our approach employs a mix of multi-iteration pre-planning and dynamic feedback techniques which are generally applicable to any recursive query implementation in an RDBMS. We also validated the effectiveness of our proposed techniques by conducting experiments on a prototype implementation using a real-life social network data from the FriendFeed online blogging service.
Ahmad Ghazal, Dawit Yimam Seid, Alain Crolotte, Mohammed Al-Kateb
SIGMOD Conference4
2010 Stratified Reservoir Sampling over Heterogeneous Data Streams
Mohammed Al-Kateb, Byung Suk Lee 0001
SSDBM1
2007 Adaptive-Size Reservoir Sampling over Data Streams
abstract
Reservoir sampling is a well-known technique for sequential random sampling over data streams. Conventional reservoir sampling assumes a fixed-size reservoir. There are situations, however, in which it is necessary and/or advantageous to adaptively adjust the size of a reservoir in the middle of sampling due to changes in data characteristics and/or application behavior. This paper studies adaptive size reservoir sampling over data streams considering two main factors: reservoir size and sample uniformity. First, the paper conducts a theoretical study on the effects of adjusting the size of a reservoir while sampling is in progress. The theoretical results show that such an adjustment may bring a negative impact on the probability of the sample being uniform (called uniformity confidence herein). Second, the paper presents a novel algorithm for maintaining the reservoir sample after the reservoir size is adjusted such that the resulting uniformity confidence exceeds a given threshold. Third, the paper extends the proposed algorithm to an adaptive multi-reservoir sampling algorithm for a practical application in which samples are collected from memory-limited wireless sensor networks using a mobile sink. Finally, the paper empirically examines the adaptivity of the multi-reservoir sampling algorithm with regard to reservoir size and sample uniformity using real sensor networks data sets.
Mohammed Al-Kateb, Byung Suk Lee 0001, Xiaoyang Sean Wang
SSDBM1
2007 Reservoir Sampling over Memory-Limited Stream Joins
abstract
In stream join processing with limited memory, uniform random sampling is useful for approximate query evaluation. In this paper, we address the problem of reservoir sampling over memory-limited stream joins. We present two sampling algorithms, reservoir join-sampling (RJS) and progressive reservoir join-sampling (PRJS). RJS is designed straightforwardly by using a fixed-size reservoir sampling on a join-sample (i.e., random sample of a join output stream). Anytime the sample in the reservoir is used, RJS always gives a uniform random sample of the original join output stream. With limited memory, however, the available memory may not be large enough even for the join buffer, thereby severely limiting the reservoir size. PRJS alleviates this problem by increasing the reservoir size during the join-sampling. This increasing is possible since the memory requirement by the join-sampling algorithm decreases over time. A larger reservoir provides a closer representation of the original join output stream. However, it comes with a negative impact on the probability of the sample being uniform. Through experiments we examine the tradeoffs and compare the two algorithms in terms of the aggregation error on the reservoir sample.
Mohammed Al-Kateb, Byung Suk Lee 0001, Xiaoyang Sean Wang
SSDBM1
2005 CME: A Temporal Relational Model for Efficient Coalescing
abstract
Coalescing is a data restructuring operation applicable to temporal databases. It merges timestamps of adjacent or overlapping tuples that have identical attribute values. The likelihood that a temporal query employs coalescing is very high. However, coalescing is an expensive and time consuming operation. In this paper, we present a novel temporal relational model through which coalescing becomes quite simple. The basic idea is to augment each time-varying attribute in a temporal relation with two additional attributes that trace changes in values of the corresponding time-varying attribute. One attribute traces changes in values with respect to each individual instance (i.e. tuples having the same key value), while the other attribute traces changes in values globally for all instances (i.e. all tuples in the temporal relation). Using these tracing attributes, coalescing could be easily implemented through a quite simple join-free group-by query. The coalescing query is fully processed and optimized by the underlying database management system.
Mohammed Al-Kateb, Essam Mansour 0001, Mohamed E. El-Sharkawi
TIME1