Amr El-Helw

dblp:80/951 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
0since 2021 · last 2018
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 5 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
7 papers
Query processing and optimization · 56% Database system architecture and tuning · 24% Distributed and cloud data management · 14%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 64% Parallel and multicore computing · 36%

Topics — the 11 heaviest of 15, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Distributed and cloud data management › federated database
federated query processing
0.312018
F1 Query: Declarative Querying at Scale · Proc. VLDB Endow. 2018
Database system architecture and tuning › parallel database system
massively parallel processing database
0.212015
Optimization of Common Table Expressions in MPP Database Systems · Proc. VLDB Endow. 2015
Query processing and optimization › runtime optimization › data skipping
partition pruning
0.212014
Optimizing queries over partitioned tables in MPP systems · SIGMOD Conference 2014
Query processing and optimization › query optimization
query optimizer architecture
0.212014
Orca: a modular query optimizer architecture for big data · SIGMOD Conference 2014
Query processing and optimization
cardinality estimation
0.222009
StatAdvisor: Recommending Statistical Views · Proc. VLDB Endow. 2009
Collecting and Maintaining Just-in-Time Statistics · ICDE 2007
Data mining › text mining
information extraction
0.112012
Just-in-time information extraction using extraction views · SIGMOD Conference 2012
Query processing and optimization › query optimization
statistics management
0.112009
StatAdvisor: Recommending Statistical Views · Proc. VLDB Endow. 2009
Query processing and optimization › query optimization › statistics management
statistics collection
0.112007
Collecting and Maintaining Just-in-Time Statistics · ICDE 2007
Query processing and optimization › query planning
query plan generation
0.112015
Optimization of Common Table Expressions in MPP Database Systems · Proc. VLDB Endow. 2015
Query processing and optimization
analytical query processing
0.112014
Orca: a modular query optimizer architecture for big data · SIGMOD Conference 2014
Query processing and optimization
query compilation
0.012007
Collecting and Maintaining Just-in-Time Statistics · ICDE 2007

Methods — techniques the papers use, named apart from their topics

declarative querying · 0.7SQL · 0.7runtime plan decisions · 0.4partition pruning · 0.4just-in-time extraction · 0.1plan-based candidate enumeration · 0.1benefit-based analysis · 0.1sensitivity analysis · 0.1
YearPublicationVenuePosition
2018 Analytical framework for adaptive compressive sensing for target detection within wireless visual sensor networks
Salema F. Fayed, Sherin M. Youssef, Amr El-Helw, Mohammad N. Patwary, Mansour Moniri
Multim. Tools Appl.3
2018 F1 Query: Declarative Querying at Scale
abstract
F1 Query is a stand-alone, federated query processing platform that executes SQL queries against data stored in different file-based formats as well as different storage systems at Google (e.g., Bigtable, Spanner, Google Spreadsheets, etc.). F1 Query eliminates the need to maintain the traditional distinction between different types of data processing workloads by simultaneously supporting: (i) OLTP-style point queries that affect only a few records; (ii) low-latency OLAP querying of large amounts of data; and (iii) large ETL pipelines. F1 Query has also significantly reduced the need for developing hard-coded data processing pipelines by enabling declarative queries integrated with custom business logic. F1 Query satisfies key requirements that are highly desirable within Google: (i) it provides a unified view over data that is fragmented and distributed over multiple data sources; (ii) it leverages datacenter resources for performant query processing with high throughput and low latency; (iii) it provides high scalability for large data sizes by increasing computational parallelism; and (iv) it is extensible and uses innovative approaches to integrate complex business logic in declarative query processing. This paper presents the end-to-end design of F1 Query. Evolved out of F1, the distributed database originally built to manage Google's advertising data, F1 Query has been in production for multiple years at Google and serves the querying needs of a large number of users and systems.
Bart Samwel, John Cieslewicz, Ben Handy, Jason Govig, Petros Venetis, Chanjun Yang, Keith Peters, Jeff Shute, Daniel Tenedorio, Himani Apte, Felix Weigel, David Wilhite, Jiexing Li, Zhan Yuan, Craig Chasseur, Ian Rae, Anurag Biyani, Andrew Harn, Andrey Gubichev, Amr El-Helw, Orri Erling, Zhepeng Yan, Mohan Yang, Yiqun Wei, Thanh Do, Colin Zheng, Goetz Graefe, Somayeh Sardashti, Ahmed M. Aly, Divyakant Agrawal, Shivakumar Venkataraman
Proc. VLDB Endow.24
2016 Adaptive compressive sensing for target tracking within wireless visual sensor networks-based surveillance applications
Salema F. Fayed, Sherin M. Youssef, Amr El-Helw, Mohammad N. Patwary, Mansour Moniri
Multim. Tools Appl.3
2015 Optimization of Common Table Expressions in MPP Database Systems
abstract
Big Data analytics often include complex queries with similar or identical expressions, usually referred to as Common Table Expressions (CTEs). CTEs may be explicitly defined by users to simplify query formulations, or implicitly included in queries generated by business intelligence tools, financial applications and decision support systems. In Massively Parallel Processing (MPP) database systems, CTEs pose new challenges due to the distributed nature of query processing, the overwhelming volume of underlying data and the scalability criteria that systems are required to meet. In these settings, the effective optimization and efficient execution of CTEs are crucial for the timely processing of analytical queries over Big Data. In this paper, we present a comprehensive framework for the representation, optimization and execution of CTEs in the context ofOrca-- Pivotal's query optimizer for Big Data. We demonstrate experimentally the benefits of our techniques using industry standard decision support benchmark.
Amr El-Helw, Venkatesh Raghavan, Mohamed A. Soliman, George C. Caragea, Zhongxian Gu, Michalis Petropoulos
Proc. VLDB Endow.1
2014 Optimizing queries over partitioned tables in MPP systems
abstract
Partitioning of tables based on value ranges provides a powerful mechanism to organize tables in database systems. In the context of data warehousing and large-scale data analysis partitioned tables are of particular interest as the nature of queries favors scanning large swaths of data. In this scenario, eliminating partitions from a query plan that contain data not relevant to answering a given query can represent substantial performance improvements. Dealing with partitioned tables in query optimization has attracted significant attention recently, yet, a number of challenges unique to Massively Parallel Processing (MPP) databases and their distributed nature remain unresolved. In this paper, we present optimization techniques for queries over partitioned tables as implemented in Pivotal Greenplum Database. We present a concise and unified representation for partitioned tables and devise optimization techniques to generate query plans that can defer decisions on accessing certain partitions to query run-time. We demonstrate, the resulting query plans distinctly outperform conventional query plans in a variety of scenarios.
Lyublena Antova, Amr El-Helw, Mohamed A. Soliman, Zhongxian Gu, Michalis Petropoulos, F. Michael Waas
SIGMOD Conference2
2014 Orca: a modular query optimizer architecture for big data
abstract
The performance of analytical query processing in data management systems depends primarily on the capabilities of the system's query optimizer. Increased data volumes and heightened interest in processing complex analytical queries have prompted Pivotal to build a new query optimizer.
Mohamed A. Soliman, Lyublena Antova, Venkatesh Raghavan, Amr El-Helw, Zhongxian Gu, Entong Shen, George C. Caragea, Carlos Garcia-Alvarado, Foyzur Rahman, Michalis Petropoulos, F. Michael Waas, Sivaramakrishnan Narayanan, Konstantinos Krikellas, Rhonda Baldwin
SIGMOD Conference4
2012 Just-in-time information extraction using extraction views
abstract
demonstration Share on Just-in-time information extraction using extraction views Authors: Amr El-Helw EMC Corp., San Mateo, CA, USA EMC Corp., San Mateo, CA, USAView Profile , Mina H. Farid University of Waterloo, Waterloo, ON, Canada University of Waterloo, Waterloo, ON, CanadaView Profile , Ihab F. Ilyas View Profile Authors Info & Claims SIGMOD '12: Proceedings of the 2012 ACM SIGMOD International Conference on Management of DataMay 2012 Pages 613–616https://doi.org/10.1145/2213836.2213913Published:20 May 2012Publication History 5citation399DownloadsMetricsTotal Citations5Total Downloads399Last 12 Months3Last 6 weeks0 Get Citation AlertsNew Citation Alert added!This alert has been successfully added and will be sent to:You will be notified whenever a record that you have chosen has been cited.To manage your alert preferences, click on the button below.Manage my Alerts New Citation Alert!Please log in to your account Save to BinderSave to BinderCreate a New BinderNameCancelCreateExport CitationPublisher SiteGet Access
Amr El-Helw, Mina H. Farid, Ihab F. Ilyas
SIGMOD Conference1
2011 Column-oriented query processing for row stores
abstract
Column-oriented DBMSs have gained increasing interest due to their superior performance for analytical workloads. Prior efforts tried to determine the possibility of simulating the query processing techniques of column-oriented systems in row-oriented databases, in a hope to improve their performance, especially for OLAP and data warehousing applications. In this paper, we show that column-oriented query processing can significantly improve the performance of row-oriented DBMSs. We introduce new operators that take into account the unique characteristics of data obtained from indexes, and exploit new technologies such as flash SSDs and multi-core processors to boost the performance. We demonstrate our approach with an experimental study using a prototype built on a commercial row-oriented DBMS.
Amr El-Helw, Kenneth A. Ross, Bishwaranjan Bhattacharjee, Christian A. Lang, George A. Mihaila
DOLAP1
2009 StatAdvisor: Recommending Statistical Views
abstract
Database statistics are crucial to cost-based optimizers for estimating the execution cost of a query plan. Using traditional basic statistics on base tables requires adopting unrealistic assumptions to estimate the cardinalities of intermediate results, which usually causes large estimation errors that can be several orders of magnitude. Modern commercial database systems support statistical or sample views, which give more accurate statistics on intermediate results and query sub-expressions. While previous research focused on creating and maintaining these advanced statistics, only little effort has been done towards automatically recommending the most beneficial statistical views to construct. In this paper, we present StatAdvisor , a system for recommending statistical views for a given SQL workload. The StatAdvisor addresses the special characteristics of statistical views with respect to view matching and benefit estimation, and introduces a novel plan-based candidate enumeration method, and a benefit-based analysis to determine the most useful statistical views. We present the basic concepts, architecture, and key features of StatAdvisor , and demonstrate its validity and benefits through an extensive experimental study using a prototype that we built in the IBM® DB2® database system as part of the DB2 Design Advisor tools.
Amr El-Helw, Ihab F. Ilyas, Calisto Zuzarte
Proc. VLDB Endow.1
2007 Collecting and Maintaining Just-in-Time Statistics
abstract
Traditional DBMSs decouple statistics collection and query optimization both in space and time. Decoupling in time may lead to outdated statistics. Decoupling in space may cause statistics not to be available at the desired granularity needed to optimize a particular query, or some important statistics may not be available at all. Overall, this decoupling often leads to large cardinality estimation errors and, in consequence, to the selection of suboptimal plans for query execution. In this paper, we present JITS, a system for proactively collecting query-specific statistics during query compilation. The system employs a lightweight sensitivity analysis to choose which statistics to collect by making use of previously collected statistics and database activity patterns. The collected statistics are materialized and incrementally updated for later reuse. We present the basic concepts, architecture, and key features of JITS. We demonstrate its benefits through an extensive experimental study on a prototype inside the IBM DB2 engine.
Amr El-Helw, Ihab F. Ilyas, Wing Lau, Volker Markl, Calisto Zuzarte
ICDE1