EDBT 2026 Demo / reviewers in the wild / expert
Ahmad Ghazal
dblp:16/5409
· DBLP profile ↗
20ranked-venue papers
11as first author
3since 2021 · last 2025
0009-0002-3071-957XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 18 · 11 first-author · 3 since 2021Artificial intelligence and machine learning · 8 · 6 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hyper: Hybrid Physical Design Advisor with Multi-agent Reinforcement LearningabstractVarious physical design (PD) options within a single database have emerged to optimize diverse workloads, including row-based PDs (e.g., index) and column-based PDs (e.g., column-store replica), each with its own acceleration advantages for different workloads. Determining the optimal combination of these two PDs is a labor-intensive and challenging task, yet it could result in significant performance improvements for the system. Recent automated index advisors (AIAs) have concentrated on identifying the most advantageous combination of row-based PDs. However, the extension of these efforts to the present problem has proven challenging due to 1) the larger search space of hybrid PD selections, 2) the inadequate consideration of the complex interactions between heterogeneous PDs, and 3) the inaccurate evaluation made by the what-if optimizer. To address these issues, we propose a Hybrid physical design advisor (Hyper) with multi-agent reinforcement learning. Hyper excels at recommending the optimal combination of PDs under any specific workload, with an overarching emphasis on both efficiency and quality. Comprehensive evaluations on well-established benchmarks show that our approach outperforms state-of-the-art methods. Yuanjia Zhang, Chengcheng Yang, Ahmad Ghazal, Rong Zhang 0002, Huiqi Hu, Xiaoju Wu, Xuan Zhou 0001 |
ICDE | 4 |
| 2024 | OSSInsight: Scalable GitHub AnalysisabstractGitHub is a platform hosting code, enabling collaboration, and supporting version control for a global community of over 100 million developers. The need for free tools is crucial for researching open-source software. Based on our research, we found out that existing tools lack real-time GitHub data processing or have limited functionalities. This demonstration presents OSSInsight, an open source tool for researching and analyzing GitHub repositories. We first present the architecture of the tool including its access to nearly 7 billion archived & real time data and how it is powered by TiDB. The demonstration shows how OSSInsight provides analysis of GitHub data along three dimensions: developers, repositories and organizations. All these analysis are based on generated SQL queries submitted to TiDB database. TiDB possesses HTAP capabilities, utilizing its row store for simple SQL queries while relying on its column store for more complex queries. Users can view and edit these SQL queries and also view their execution plan. Finally, OSSInsight provides an innovative tool based on OpenAI, that conducts data analysis using input in English text, yielding visual representations in the form of charts and graphs. Ahmad Ghazal, Zhiyuan Liang, Sunny Bains, Hanumath Maduri |
Proc. VLDB Endow. | 1 |
| 2021 | DBSpinner: Making a Case for Iterative Processing in DatabasesabstractRelational database management systems (RDBMS) have limited iterative processing support. Recursive queries were added to ANSI SQL, however, their semantics do not allow aggregation functions, which disqualifies their use for several applications, such as PageRank and shortest path computations. Recently, another SQL extension, iterative Common Table Expressions (CTEs), is proposed to enable users to perform general iterative computations on RDBMSs.In this work1, we demonstrate how iterative CTEs can be efficiently incorporated into a production RDBMS without major intrusion to the system. We have prototyped our approach on Futurewei's MPPDB, a shared nothing relational parallel database engine. The implementation is based on a functional rewrite that translates iterative CTEs to other existing SQL operators. Thus, query plans of iterative CTEs can be optimized and executed by the engine with minimal modification to the code base. We have also applied several optimizations specifically for iterative CTEs to i) minimize data movement, ii) reuse results that remain constant and iii) push down predicates to avoid unnecessary data processing. We verified our implementation through extensive experimental evaluation using real world datasets and queries. The results show the feasibility of the rewrite approach and the effectiveness of the optimizations, which improve performance by an order of magnitude in some cases. Sofoklis Floratos, Ahmad Ghazal, Jason Sun, Jianjun Chen 0001, Xiaodong Zhang 0001 |
ICDE | 2 |
| 2019 | Data Management at Huawei: Recent Accomplishments and Future ChallengesabstractHuawei is a leading global provider of information and communication technologies (ICT) infrastructure and smart devices. With integrated solutions across four key domains: telecommunication networks, IT, smart devices, and cloud services, Huawei is committed to bringing digital transformation to every person, home and organization for a fully connected and intelligent world. Founded in 1987, Huawei currently has more than 180,000 employees, and operates in more than 170 countries and regions with revenue over 100 billion USD in 2018. Data management plays a key role in all of the four key domains above. We have developed innovative products and solutions to support rapid business growth driven by customer requirements. While many data management problems are common, each domain also has its own special requirements and challenges. In this paper, we will go through recent advancements in Huawei data management technologies including a petabyte scale enterprise analytics platform (FusionInsight MPPDB) and a highly available in-memory database for telecommunication networks (GMDB). In addition, we discuss data management challenges that we are facing in the areas of autonomous databases and device-edge-cloud collaboration data platforms. Jianjun Chen 0001, Zhibiao Chen, Ahmad Ghazal, Guoliang Li 0001, Sihao Li, Weijie Ou, Mingyi Zhang 0001, Minqi Zhou |
ICDE | 4 |
| 2018 | FusionInsight LibrA: Huawei's Enterprise Cloud Data Analytics PlatformabstractHuawei Fusion Insight Libr A (FI-MPPDB) is a petabyte scale enterprise analytics platform developed by the Huawei data-base group. It started as a prototype more than five years ago, and is now being used by many enterprise customers over the globe, including some of the world's largest financial institutions. Our product direction and enhancements have been mainly driven by customer requirements in the fast evolving Chinese market. This paper describes the architecture of FI-MPPDB and some of its major enhancements. In particular, we focus on top four requirements from our customers related to data analytics on the cloud: system availability, auto tuning, query over heterogeneous data models on the cloud, and the ability to utilize powerful modern hardware for good performance. We present our latest advancements in the above areas including online expansion, auto tuning in query optimizer, SQL on HDFS, and intelligent JIT compiled execution. Finally, we present some experimental results to demonstrate the effectiveness of these technologies. Le Cai, Jianjun Chen 0001, Kuorong Chiang, Marko A. Dimitrijevic, Yonghua Ding, Ahmad Ghazal, Jacques Hebert, Kamini Jagtiani, Suzhen Lin, Demai Ni, Chunfeng Pei, Jason Sun, Li Zhang 0132, Mingyi Zhang 0001 |
Proc. VLDB Endow. | 9 |
| 2017 | BigBench V2: The New and Improved BigBenchabstractBenchmarking Big Data solutions has been gaining a lot of attention from research and industry. BigBench is one of the most popular benchmarks in this area which was adopted by the TPC as TPCx-BB. BigBench, however, has key shortcomings. The structured component of the data model is the same as the TPC-DS data model which is a complex snowflake-like schema. This is contrary to the simple star schema Big Data models in real life. BigBench also treats the semi-structured web-logs more or less as a structured table. In real life, web-logs are modeled as key-value pairs with unknown schema. Specific keys are captured at query time - a process referred to as late binding. In addition, eleven (out of thirty) of the BigBench queries are TPC-DS queries. These queries are complex SQL applied on the structured part of the data model which again is not typical of Big Data workloads. In this paper1, we present BigBench V2 to address the aforementioned limitations of the original BigBench. BigBench V2 is completely independent of TPC-DS with a new data model and an overhauled workload. The new data model has a simple structured data model. Web-logs are modeled as key-value pairs with a substantial and variable number of keys. BigBench V2 mandates late binding by requiring query processing to be done directly on key-value web-logs rather than a pre-parsed form of it. A new scale factor-based data generator is implemented to produce structured tables, key-value semistructured web-logs, and unstructured data. We implemented and executed BigBench V2 on Hive. Our proof of concept shows the feasibility of BigBench V2 and outlines different ways of implementing late binding. Ahmad Ghazal, Todor Ivanov, Pekka Kostamaa, Alain Crolotte, Ryan Voong, Mohammed Al-Kateb, Waleed Ghazal, Roberto V. Zicari |
ICDE | 1 |
| 2013 | Temporal query processing in TeradataabstractThe importance of temporal data management is evident by the temporal features recently released in major commercial database systems. In Teradata, the temporal feature is based on the TSQL2 specification. In this paper, we present Teradata's implementation approach for temporal query processing. There are two common approaches to support temporal query processing in a database engine. One is through functional query rewrites to convert a temporal query to a semantically-equivalent non-temporal counterpart, mostly by adding time-based constraints. The other is a native support that implements temporal database operations such as scans and joins directly in the DBMS internals. These approaches have competing pros and cons. The rewrite approach is generally simpler to implement. But it adds a structural complexity to original query, which can pose a potential challenge to query optimizer and cause it to generate sub-optimal plans. A native support is expected to perform better. But it usually involves a higher cost of implementation, maintenance, and extension. We discuss why and describe how Teradata adopted the rewrite approach. In addition, we present an evaluation of our approach through a performance study conducted on a variation of the TPC-H benchmark with temporal tables and queries. Mohammed Al-Kateb, Ahmad Ghazal, Alain Crolotte, Ramesh Bhashyam, Jaiprakash Chimanchode, Sai Pavan Pakala |
EDBT | 2 |
| 2013 | BigBench: towards an industry standard benchmark for big data analyticsabstractThere is a tremendous interest in big data by academia, industry and a large user base. Several commercial and open source providers unleashed a variety of products to support big data storage and processing. As these products mature, there is a need to evaluate and compare the performance of these systems. Ahmad Ghazal, Tilmann Rabl, Minqing Hu, Francois Raab, Meikel Pöss, Alain Crolotte, Hans-Arno Jacobsen |
SIGMOD Conference | 1 |
| 2012 | A new tool for multi-level partitioning in teradataabstractThis paper introduces a new tool that recommends an optimized partitioning solution called Multi-Level Partitioned Primary Index (MLPPI) for a fact table based on the queries in the workload. The tool implements a new technique using a greedy algorithm for search space enumeration. The space is driven by predicates in the queries. This technique fits very well the Teradata MLPPI scheme, as it is based on a general framework using general expressions, ranges and case expressions for partition definitions. The cost model implemented in the tool is based on the Teradata optimizer, and it is used to prune the search space for reaching a final solution. The tool resides completely on the client, and interfaces the database through APIs as opposed to previous work that requires optimizer code extension. The APIs are used to simplify the workload queries, and to capture fact table predicates and costs necessary to make the recommendation. The predicate-driven method implemented by the tool is general, and it can be applied to any clustering or partitioning scheme based on simple field expressions or complex SQL predicates. Experimental results given a particular workload will show that the recommendation from the tool outperforms a human expert. The experiments also show that the solution is scalable both with the workload complexity and the size of the fact table. Young-Kyoon Suh, Ahmad Ghazal, Alain Crolotte, Pekka Kostamaa |
CIKM | 2 |
| 2012 | The F&A Methodology and Its Experimental Validation on a Real-Life Parallel Processing Database SystemabstractThis paper complements our previous results in the context of effectively and efficient designing Parallel Relational Data Warehouses (PRDW) over heterogeneous database clusters, which are represented by the proposal of a methodology called Fragmentation & Allocation (F& A). The main merit of F& A is that of combining the fragmentation and the allocation phases simultaneously, which are instead performed separately by traditional approaches. In this paper, we prove the practical impact and the reliability of F& A on a real-life parallel processing database system. Ladjel Bellatreche, Soumia Benkrid, Alain Crolotte, Alfredo Cuzzocrea, Ahmad Ghazal |
CISIS | 5 |
| 2012 | An Efficient SQL Rewrite Approach for Temporal Coalescing in the Teradata RDBMS
Mohammed Al-Kateb, Ahmad Ghazal, Alain Crolotte |
DEXA (2) | 2 |
| 2012 | Adaptive optimizations of recursive queries in teradataabstractRecursive queries were introduced as part of ANSI SQL 99 to support processing of hierarchical data typical of air flight schedules, bill-of-materials, data cube dimension hierarchies, and ancestor-descendant information (e.g. XML data stored in relations). Recently, recursive queries have also found extensive use in web data analysis such as social network and click stream data. Teradata implemented recursive queries in V2R6 using static plans whereby a query is executed in multiple iterations, each iteration corresponding to one level of the recursion. Such a static planning strategy may not be optimal since the demographics of intermediate results from recursive iterations often vary to a great extent. Gathering feedback at each iteration could address this problem by providing size estimates to the optimizer which, in turn, can produce an execution plan for the next iteration. However, such a full feedback scheme suffers from lack of pipelining and the inability to exploit global optimizations across the different recursion iterations. In this paper, we propose adaptive optimization techniques that avoid the issues with static as well as full feedback optimization approaches. Our approach employs a mix of multi-iteration pre-planning and dynamic feedback techniques which are generally applicable to any recursive query implementation in an RDBMS. We also validated the effectiveness of our proposed techniques by conducting experiments on a prototype implementation using a real-life social network data from the FriendFeed online blogging service. Ahmad Ghazal, Dawit Yimam Seid, Alain Crolotte, Mohammed Al-Kateb |
SIGMOD Conference | 1 |
| 2011 | Verification of Partitioning and Allocation Techniques on Teradata DBMS
Ladjel Bellatreche, Soumia Benkrid, Ahmad Ghazal, Alain Crolotte, Alfredo Cuzzocrea |
ICA3PP (1) | 3 |
| 2010 | Merging Views Containing Outer Joins in the Teradata DBMS
Ahmad Ghazal, Dawit Yimam Seid, Alain Crolotte |
DEXA (1) | 1 |
| 2009 | Dynamic plan generation for parameterized queriesabstractQuery processing in a DBMS typically involves two distinct phases: compilation, which generates the best plan and its corresponding execution steps, and execution, which evaluates these steps against database objects. For some queries, considerable resource savings can be achieved by skipping the compilation phase when the same query was previously submitted and its plan was already cached. In a number of important applications the same query, called a Parameterized Query (PQ), is repeatedly submitted in the same basic form but with different parameter values. PQ's are extensively used in both data update (e.g. batch update programs) and data access queries. There are tradeoffs associated with caching and re-using query plans such as space utilization and maintenance cost. Besides, pre-compiled plans may be suboptimal for a particular execution due to various reasons including data skew and inability to exploit value-based query transformation like materialized view rewrite and unsatisfiable predicate elimination. We address these tradeoffs by distinguishing two types of plans for PQ's: generic and specific plans. Generic plans are pre-compiled plans that are independent of the actual parameter values. Prior to execution, parameter values are plugged in to generic plans. In specific plans, parameter values are plugged prior to the compilation phase. This paper provides a practical framework for dynamically deciding between specific and generic plans for PQ's based on a mix of rule and cost based heuristics which are implemented in the Teradata 12.0 DBMS. Ahmad Ghazal, Dawit Yimam Seid, Ramesh Bhashyam, Alain Crolotte, Manjula Koppuravuri, Vinod G |
SIGMOD Conference | 1 |
| 2008 | Exploiting Interactions among Query Rewrite Rules in the Teradata DBMS
Ahmad Ghazal, Dawit Yimam Seid, Alain Crolotte, Bill McKenna |
DEXA | 1 |
| 2006 | Recursive SQL Query Optimization with k-Iteration Lookahead
Ahmad Ghazal, Alain Crolotte, Dawit Yimam Seid |
DEXA | 1 |
| 2004 | Outer Join Elimination in the Teradata RDBMS
Ahmad Ghazal, Alain Crolotte, Ramesh Bhashyam |
DEXA | 1 |
| 2003 | Block Optimization in the Teradata RDBMS
Ahmad Ghazal, Ramesh Bhashyam, Alain Crolotte |
DEXA | 1 |
| 2001 | Dynamic Constraints Derivation and Maintenance in the Teradata RDBMS
Ahmad Ghazal, Ramesh Bhashyam |
DEXA | 1 |