EDBT 2026 Demo / reviewers in the wild / expert
F. Michael Waas
dblp:w/FlorianWaas · also Florian M. Waas, Florian Waas
· DBLP profile ↗
21ranked-venue papers
2as first author
3since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 19 · 2 first-author · 3 since 2021Artificial intelligence and machine learning · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | How Global Retailer ADEO Migrated to Google BigQuery with Database VirtualizationabstractWe describe how multi-national retailer ADEO successfully employed database virtualization to migrate all workloads of a complex Enterprise Data Warehouse (EDW) from a legacy Teradata system to Google BigQuery. We demonstrate the generality of the technology and the approach. Ehab Abdelhamid, Amirhossein Aleyasen, Michael Duller, Eric Foratier, Vincent Fruleux, Mirella Katch, Gourab Mitra, Rima Mutreja, Jozsef Patvarczki, Matthew Pope, Nikos Tsikoudis, F. Michael Waas |
IEEE Big Data | 12 |
| 2023 | Adaptive Real-time Virtualization of Legacy ETL Pipelines in Cloud Data Warehouses
Ehab Abdelhamid, Nikos Tsikoudis, Michael Duller, Marc Sugiyama, Nicholas E. Marino, F. Michael Waas |
EDBT | 6 |
| 2022 | Intelligent Automated Workload Analysis for Database ReplatformingabstractPerforming a detailed workload analysis is a crucial step in determining the feasibility, timeline and cost of a major data warehouse replatforming project, i.e., migration from one platform to another. A large company's data warehouse applications may include millions of queries, some of which will use features that are unsupported or have different semantics in the new warehouse, or may have poor performance there. Amirhossein Aleyasen, Mark Morcos, Lyublena Antova, Marc Sugiyama, Dmitri Korablev, Jozsef Patvarczki, Rima Mutreja, Michael Duller, F. Michael Waas, Marianne Winslett |
SIGMOD Conference | 9 |
| 2020 | A Framework for Emulating Database Operations in Cloud Data WarehousesabstractIn recent years, increased interest in cloud-based data warehousing technologies has emerged with many enterprises moving away from on-premise data warehousing solutions. The incentives for adopting cloud data warehousing technologies are many: cost-cutting, on-demand pricing, offloading data centers, unlimited hardware resources, built-in disaster recovery, to name a few. There is inherent difference in the language surface and feature sets of on-premise and cloud data warehousing solutions. This could range from subtle syntactic and semantic differences, with potentially big impact on result correctness, to complete features that exist in one system but are missing in other systems. While there have been some efforts to help automate the migration of on-premise applications to new cloud environments, a major challenge that slows down the migration pace is the handling of features not yet supported, or partially supported, by the cloud technologies. In this paper we build on our earlier work in adaptive data virtualization and present novel techniques that allow running applications utilizing sophisticated database features within foreign query engines lacking the native support of such features. In particular, we introduce a framework to manage discrepancy of metadata across heterogeneous query engines, and various mechanisms to emulate database applications code in cloud environments without any need to rewrite or change the application code. Mohamed A. Soliman, Lyublena Antova, Marc Sugiyama, Michael Duller, Amirhossein Aleyasen, Gourab Mitra, Ehab Abdelhamid, Mark Morcos, Michele Gage, Dmitri Korablev, F. Michael Waas |
SIGMOD Conference | 11 |
| 2018 | High-Throughput Adaptive Data Virtualization via Context-Aware Query RoutingabstractQuery throughput degradation is a common problem when migrating applications from an on-premises to a cloud-based data warehouse. In this paper, we propose CAQR: a context-aware query router. CAQR is a middleware-based database replication solution that improves workload throughput by combining the strong points of existing data replication methods and mitigating their limitations at the same time. CAQR provides strong data consistency without the need for real-time synchronization or manual annotation of queries. Employing a master-slave architecture, CAQR runs atop an adaptive data virtualization layer and automatically sends a query to the correct replica, so that the execution preserves the strong data consistency property. Extensive experiments on real-world and benchmark workloads show the practicality and efficiency of the solution: CAQR achieves near-linear scalability, without requiring any changes to the application layer. Amirhossein Aleyasen, Mohamed A. Soliman, Lyublena Antova, F. Michael Waas, Marianne Winslett |
IEEE BigData | 4 |
| 2018 | Rapid Adoption of Cloud Data Warehouse Technology Using Datometry Hyper-QabstractThe database industry is about to undergo a fundamental transformation of unprecedented magnitude as enterprises start trading their well-established database stacks on premises for cloud database technology in order to take advantage of the economics cloud service providers have long promised. Industry experts and analysts expect the next years to prove a watershed moment in this transformation, as cloud databases finally reached critical mass and maturity. Lyublena Antova, Derrick Bryant, Tuan Cao, Michael Duller, Mohamed A. Soliman, F. Michael Waas |
SIGMOD Conference | 6 |
| 2016 | Datometry Hyper-Q: Bridging the Gap Between Real-Time and Historical AnalyticsabstractWall Street's trading engines are complex database applications written for time series databases like kdb+ that uses the query language Q to perform real-time analysis. Extending the models to include other data sources, e.g., historic data, is critical for backtesting and compliance. However, Q applications cannot run directly on SQL databases. Therefore, financial institutions face the dilemma of either maintaining two separate application stacks, one written in Q and the other in SQL, which means increased IT cost and increased risk, or migrating all Q applications to SQL, which results in losing the inherent competitive advantage on Q real-time processing. Neither solution is desirable as both alternatives are costly, disruptive, and suboptimal. In this paper we present Hyper-Q, a data virtualization plat- form that overcomes the chasm. Hyper-Q enables Q applications to run natively on PostgreSQL-compatible databases by translating queries and results on the fly. We outline the basic concepts, detail specific difficulties, and demonstrate the viability of the approach with a case study. Lyublena Antova, Rhonda Baldwin, Derrick Bryant, Tuan Cao, Michael Duller, John Eshleman, Zhongxian Gu, Entong Shen, Mohamed A. Soliman, F. Michael Waas |
SIGMOD Conference | 10 |
| 2014 | Optimizing queries over partitioned tables in MPP systemsabstractPartitioning of tables based on value ranges provides a powerful mechanism to organize tables in database systems. In the context of data warehousing and large-scale data analysis partitioned tables are of particular interest as the nature of queries favors scanning large swaths of data. In this scenario, eliminating partitions from a query plan that contain data not relevant to answering a given query can represent substantial performance improvements. Dealing with partitioned tables in query optimization has attracted significant attention recently, yet, a number of challenges unique to Massively Parallel Processing (MPP) databases and their distributed nature remain unresolved. In this paper, we present optimization techniques for queries over partitioned tables as implemented in Pivotal Greenplum Database. We present a concise and unified representation for partitioned tables and devise optimization techniques to generate query plans that can defer decisions on accessing certain partitions to query run-time. We demonstrate, the resulting query plans distinctly outperform conventional query plans in a variety of scenarios. Lyublena Antova, Amr El-Helw, Mohamed A. Soliman, Zhongxian Gu, Michalis Petropoulos, F. Michael Waas |
SIGMOD Conference | 6 |
| 2014 | Orca: a modular query optimizer architecture for big dataabstractThe performance of analytical query processing in data management systems depends primarily on the capabilities of the system's query optimizer. Increased data volumes and heightened interest in processing complex analytical queries have prompted Pivotal to build a new query optimizer. Mohamed A. Soliman, Lyublena Antova, Venkatesh Raghavan, Amr El-Helw, Zhongxian Gu, Entong Shen, George C. Caragea, Carlos Garcia-Alvarado, Foyzur Rahman, Michalis Petropoulos, F. Michael Waas, Sivaramakrishnan Narayanan, Konstantinos Krikellas, Rhonda Baldwin |
SIGMOD Conference | 11 |
| 2011 | Dynamic prioritization of database queriesabstractEnterprise database systems handle a variety of diverse query workloads that are of different importance to the business. For example, periodic reporting queries are usually mission critical whereas ad-hoc queries by analysts tend to be less crucial. It is desirable to enable database administrators to express (and modify) the importance of queries at a simple and intuitive level. The mechanism used to enforce these priorities must be robust, adaptive and efficient. In this paper, we present a mechanism that continuously determines and re-computes the ideal target velocity of concurrent database processes based on their run-time statistics to achieve this prioritization. In this scheme, every process autonomously adjusts its resource consumption using basic control theory principles. The self-regulating and decentralized design of the system enables effective prioritization even in the presence of exceptional situations, including software defects or unexpected/unplanned query termination with no measurable overhead. We have implemented this approach in Greenplum Parallel Database and demonstrate its effectiveness and general applicability in a series of experiments. Sivaramakrishnan Narayanan, F. Michael Waas |
ICDE | 2 |
| 2011 | Online Expansion of Largescale Data Warehouses
Jeffrey Cohen, John Eshleman, Brian Hagenbuch, Joy Kent, Christopher Pedrotti, Gavin Sherry, F. Michael Waas |
Proc. VLDB Endow. | 7 |
| 2010 | Database architecture (R)evolution: New hardware vs. new softwareabstractThe last few years have been exciting for data management system designers. The explosion in user and enterprise data coupled with the availability of newer, cheaper, and more capable hardware have lead system designers and researchers to rethink and, in some cases, reinvent the traditional DBMS architecture. In the space of data warehousing and analytics alone, more than a dozen new database product offerings have recently appeared, and dozens of research system papers are routinely published each year. In this panel, we will ask our panelists, a mix of industry and academic experts, which of those trends will have lasting effects on database system design, and which directions hold the biggest potential for future research. We are particularly interested in the differences in views and approaches between academic and industrial research. Stavros Harizopoulos, Tassos Argyros, Peter Boncz, Dan Dietterich, Samuel Madden 0001, F. Michael Waas |
ICDE | 6 |
| 2009 | Parallelizing extensible query optimizersabstractQuery optimization is the most computationally complex task in a database management systems. In many query optimizers, faster CPUs and increased RAM can translate directly to better query plans and thus better overall system performance. Although memory size continues to scale with Moore's Law, processor speeds are leveling off. Chip manufacturers are now focusing on multicore designs that integrate increasing numbers of cores in a single CPU. Query optimizers need to be parallelized in order to continue enjoying the growth trend of Moore's Law. F. Michael Waas, Joseph M. Hellerstein |
SIGMOD Conference | 1 |
| 2005 | Database Change Notifications: Primitives for Efficient Database Query Result Caching
César A. Galindo-Legaria, Torsten Grabs, Christian Kleinerman, F. Michael Waas |
VLDB | 4 |
| 2004 | Query Processing for SQL UpdatesabstractA rich set of concepts and techniques has been developed in the context of query processing for the efficient and robust execution of queries. So far, this work has mostly focused on issues related to data-retrieval queries, with a strong backing on relational algebra. However, update operations can also exhibit a number of query processing issues, depending on the complexity of the operations and the volume of data to process. Such issues include lookup and matching of values, navigational vs. set-oriented algorithms and trade-offs between plans that do serial or random I/Os.In this paper we present an overview of the basic techniques used to support SQL DML (Data Manipulation Language) in Microsoft SQL Server. Our focus is on the integration of update operations into the query processor, the query execution primitives required to support updates, and the update-specific considerations to analyze and execute update plans. Full integration of update processing in the query processor provides a robust and flexible framework and leverages existing query processing techniques. César A. Galindo-Legaria, Stefano Stefani, F. Michael Waas |
SIGMOD Conference | 3 |
| 2003 | Statistics on Views
César A. Galindo-Legaria, Milind Joshi, F. Michael Waas, Ming-Chuan Wu |
VLDB | 3 |
| 2002 | The Effect Of Cost Distributions On Evolutionary Optimization Algorithms
César A. Galindo-Legaria, F. Michael Waas |
GECCO | 2 |
| 2002 | XMark: A Benchmark for XML Data Management
Albrecht Schmidt 0002, F. Michael Waas, Martin L. Kersten, Michael J. Carey 0001, Ioana Manolescu, Ralph Busse |
VLDB | 2 |
| 2001 | FeedbackBypass: A New Approach to Interactive Similarity Query Processing
Ilaria Bartolini, Paolo Ciaccia, F. Michael Waas |
VLDB | 3 |
| 2000 | Counting, Enumerating, and Sampling of Execution Plans in a Cost-Based Query OptimizerabstractTesting an SQL database system by running large sets of deterministic or stochastic SQL statements is common practice in commercial database development. However, code defects often remain undetected as the query optimizer's choice of an execution plan is not only depending on the query but strongly influenced by a large number of parameters describing the database and the hardware environment. Modifying these parameters in order to steer the optimizer to select other plans is difficult since this means anticipating often complex search strategies implemented in the optimizer. F. Michael Waas, César A. Galindo-Legaria |
SIGMOD Conference | 1 |
| 1997 | Load Balanced Query Evaluation in Shared-Everything Environments
Stefan Manegold, Johann K. Obermaier, F. Michael Waas |
Euro-Par | 3 |