Andrew Witkowski

dblp:63/3671 · DBLP profile ↗
← Back
17ranked-venue papers
6as first author
0since 2021 · last 2020
0009-0001-6118-8380ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 15 · 5 first-authorSystems, architecture and hardware · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
16 papers
Query processing and optimization · 49% Data mining · 17% Database system architecture and tuning · 13%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 78% Parallel and multicore computing · 19% Hardware accelerators and domain-specific architectures · 2%

Topics — the 30 heaviest of 37, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
view maintenance
0.532020
Automated Generation of Materialized Views in Oracle · Proc. VLDB Endow. 2020
Optimizing Refresh of a Set of Materialized Views · VLDB 2005
Materialized Views in Oracle · VLDB 1998
Query processing and optimization › materialized view
materialized view selection
0.412020
Automated Generation of Materialized Views in Oracle · Proc. VLDB Endow. 2020
Database system architecture and tuning
self-managing database systems
0.412020
Automated Generation of Materialized Views in Oracle · Proc. VLDB Endow. 2020
Data mining
clustering
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Data mining › data reduction
data pruning
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Query processing and optimization › runtime optimization › data skipping
partition pruning
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Indexing and storage engines › synopsis structure
zone maps
0.312017
Dimensions Based Data Clustering and Zone Maps · Proc. VLDB Endow. 2017
Query processing and optimization › query rewriting
query transformation
0.222009
Enhanced Subquery Optimizations in Oracle · Proc. VLDB Endow. 2009
Cost-Based Query Transformation in Oracle · VLDB 2006
Data models and query languages › SQL
SQL extension
0.132005
Advanced SQL modeling in RDBMS · ACM Trans. Database Syst. 2005
Data Densification in a Relational Database System · SIGMOD Conference 2004
Spreadsheets in RDBMS for OLAP · SIGMOD Conference 2003
Cloud and datacenter computing
database-as-a-service
0.112020
Automated Generation of Materialized Views in Oracle · Proc. VLDB Endow. 2020
Query processing and optimization › query optimization
nested query optimization
0.112009
Enhanced Subquery Optimizations in Oracle · Proc. VLDB Endow. 2009
Data stream processing
continuous query processing
0.112007
Continuous Queries in Oracle · VLDB 2007
Data models and query languages
query language
0.122005
Advanced SQL modeling in RDBMS · ACM Trans. Database Syst. 2005
Query By Excel · VLDB 2005
Data models and query languages
query interface
0.122005
Query By Excel · VLDB 2005
Business Modelling Using SQL Spreadsheets · VLDB 2003
Query processing and optimization › query rewriting
cost-based query rewriting
0.112006
Cost-Based Query Transformation in Oracle · VLDB 2006
Distributed and cloud data management
array processing
0.112005
Advanced SQL modeling in RDBMS · ACM Trans. Database Syst. 2005
Query processing and optimization
query optimization
0.112005
Advanced SQL modeling in RDBMS · ACM Trans. Database Syst. 2005
Data mining
business intelligence
0.012004
Where is Business Intelligence taking today's Database Systems? · VLDB 2004
Query processing and optimization › OLAP
OLAP query optimization
0.012003
Spreadsheets in RDBMS for OLAP · SIGMOD Conference 2003
Query processing and optimization › OLAP
OLAP query processing
0.012003
Spreadsheets in RDBMS for OLAP · SIGMOD Conference 2003
Data models and query languages
SQL
0.012003
Business Modelling Using SQL Spreadsheets · VLDB 2003
Parallel and multicore computing
parallel query processing
0.012009
Enhanced Subquery Optimizations in Oracle · Proc. VLDB Endow. 2009
Query processing and optimization
query execution
0.012007
Continuous Queries in Oracle · VLDB 2007
Query processing and optimization
materialized view
0.011998
Materialized Views in Oracle · VLDB 1998
Database system architecture and tuning › analytical database system
ROLAP
0.012005
Advanced SQL modeling in RDBMS · ACM Trans. Database Syst. 2005
Database system architecture and tuning
database machine
0.011993
NCR 3700 - The Next-Generation Industrial Database Computer · VLDB 1993
Database system architecture and tuning
commercial database systems
0.011998
Materialized Views in Oracle · VLDB 1998
Transaction processing and concurrency control › concurrency control
multiversion concurrency control
0.011986
Performance Evaluation of Multiversion with the Oracle Synchronization · SIGMETRICS 1986
Transaction processing and concurrency control › concurrency control
timestamp ordering
0.011986
Performance Evaluation of Multiversion with the Oracle Synchronization · SIGMETRICS 1986
Hardware accelerators and domain-specific architectures
database accelerator
0.011993
NCR 3700 - The Next-Generation Industrial Database Computer · VLDB 1993

Methods — techniques the papers use, named apart from their topics

extended covering sub-expression algorithm · 0.9cost-based selection · 0.9window functions · 0.2antijoin · 0.2query optimization · 0.1parameterization · 0.1access structure design · 0.1partitioned outer join · 0.0array-based execution models · 0.0MOLAP engines · 0.0two-dimensional poisson process · 0.0queueing analysis · 0.0stochastic process analysis · 0.0linear poisson processes · 0.0
YearPublicationVenuePosition
2020 Automated Generation of Materialized Views in Oracle
abstract
Automated generation of a right set of materialized views is a challenging task. It is a highly desirable feature for autonomous databases. The selection of materialized views must be based on cost and verifiable in the actual database environment. This paper describes an automated system that generates, selects, verifies, and maintains materialized views in the Oracle RDBMS; it presents a novel technique, called the extended covering sub-expression algorithm, for the automated generation of materialized views. An extensive set of experiments is described that demonstrates the feasibility and efficiency of this approach. This system has been fully implemented and is going to be deployed on the Oracle Autonomous Database on the Cloud.
Rafi Ahmed, Randall G. Bello, Andrew Witkowski
Proc. VLDB Endow.3
2017 Dimensions Based Data Clustering and Zone Maps
abstract
In recent years, the data warehouse industry has witnessed decreased use of indexing but increased use of compression and clustering of data facilitating efficient data access and data pruning in the query processing area. A classic example of data pruning is the partition pruning, which is used when table data is range or list partitioned. But lately, techniques have been developed to prune data at a lower granularity than a table partition or sub-partition. A good example is the use of data pruning structure called zone map. A zone map prunes zones of data from a table on which it is defined. Data pruning via zone map is very effective when the table data is clustered by the filtering columns. The database industry has offered support to cluster data in tables by its local columns, and to define zone maps on clustering columns of such tables. This has helped improve the performance of queries that contain filter predicates on local columns. However, queries in data warehouses are typically based on star/snowflake schema with filter predicates usually on columns of the dimension tables joined to a fact table. Given this, the performance of data warehouse queries can be significantly improved if the fact table data is clustered by columns of dimension tables together with zone maps that maintain min/max value ranges of these clustering columns over zones of fact table data. In recognition of this opportunity of significantly improving the performance of data warehouse queries, Oracle 12c release 1 has introduced the support for dimension based clustering of fact tables together with data pruning of the fact tables via dimension based zone maps.
Mohamed Ziauddin, Andrew Witkowski, You Jung Kim, Janaki Lahorani, Dmitry Potapov, Murali Krishna
Proc. VLDB Endow.2
2009 Enhanced Subquery Optimizations in Oracle
abstract
This paper describes enhanced subquery optimizations in Oracle relational database system. It discusses several techniques -- subquery coalescing, subquery removal using window functions, and view elimination for group-by queries. These techniques recognize and remove redundancies in query structures and convert queries into potentially more optimal forms. The paper also discusses novel parallel execution techniques, which have general applicability and are used to improve the scalability of queries that have undergone some of these transformations. It describes a new variant of antijoin for optimizing subqueries involved in the universal quantifier with columns that may have nulls. It then presents performance results of these optimizations, which show significant execution time improvements.
Srikanth Bellamkonda, Rafi Ahmed, Andrew Witkowski, Angela Amor, Mohamed Zaït, Chun Chieh Lin
Proc. VLDB Endow.3
2007 Continuous Queries in Oracle
Sankar Subramanian, Srikanth Bellamkonda, Hua-Gang Li, Vince Liang, Wayne Smith, James Terry, Tsae-Feng Yu, Andrew Witkowski
VLDB9
2006 Cost-Based Query Transformation in Oracle
Rafi Ahmed, Allison W. Lee, Andrew Witkowski, Dinesh Das, Hong Su, Mohamed Zaït, Thierry Cruanes
VLDB3
2005 Optimizing Refresh of a Set of Materialized Views
Nathan Folkert, Abhinav Gupta 0003, Andrew Witkowski, Sankar Subramanian, Srikanth Bellamkonda, Shrikanth Shankar, Tolga Bozkaya
VLDB3
2005 Query By Excel
Andrew Witkowski, Srikanth Bellamkonda, Tolga Bozkaya, Aman Naimat, Sankar Subramanian, Allison Waingold
VLDB1
2005 Advanced SQL modeling in RDBMS
abstract
Commercial relational database systems lack support for complex business modeling. ANSI SQL cannot treat relations as multidimensional arrays and define multiple, interrelated formulas over them, operations which are needed for business modeling. Relational OLAP (ROLAP) applications have to perform such tasks using joins, SQL Window Functions, complex CASE expressions, and the GROUP BY operator simulating the pivot operation. The designated place in SQL for calculations is the SELECT clause, which is extremely limiting and forces the user to generate queries with nested views, subqueries and complex joins. Furthermore, SQL query optimizers are preoccupied with determining efficient join orders and choosing optimal access methods and largely disregard optimization of multiple, interrelated formulas. Research into execution methods has thus far concentrated on efficient computation of data cubes and cube compression rather than on access structures for random, interrow calculations. This has created a gap that has been filled by spreadsheets and specialized MOLAP engines, which are good at specification of formulas for modeling but lack the formalism of the relational model, are difficult to coordinate across large user groups, exhibit scalability problems, and require replication of data between the tool and RDBMS. This article presents an SQL extension called SQL Spreadsheet , to provide array calculations over relations for complex modeling. We present optimizations, access structures, and execution models for processing them efficiently. Special attention is paid to compile time optimization for expensive operations like aggregation. Furthermore, ANSI SQL does not provide a good separation between data and computation and hence cannot support parameterization for SQL Spreadsheets models. We propose two parameterization methods for SQL. One parameterizes ANSI SQL view using subqueries and scalars, which allows passing data to SQL Spreadsheet. Another method presents parameterization of the SQL Spreadsheet formulas. This supports building stand-alone SQL Spreadsheet libraries. These models are then subject to the SQL Spreadsheet optimizations during model invocation time.
Andrew Witkowski, Srikanth Bellamkonda, Tolga Bozkaya, Nathan Folkert, Abhinav Gupta 0003, John Haydu, Sankar Subramanian
ACM Trans. Database Syst.1
2004 Data Densification in a Relational Database System
abstract
Data in a relational data warehouse is usually sparse. That is, if no value exists for a given combination of dimension values, no row exists in the fact table. Densities of 0.1-2% are very common. However, users may want to view the data in a dense form, with rows for all combination of dimension values displayed even when no fact data exists for them. For example, if a product did not sell during a particular time period, users may still want to see the product for that time period with zero sales value next to it. Moreover, analytic window functions [1] and the SQL model clause [2] can more easily express time series calculations if data is dense along the time dimension because dense data will fill a consistent number of rows for each period.Data densification is the process of converting spare data into dense form. The current SQL technique for densification (using the combination of DISTINCT, CROSS JOIN and OUTER JOIN operations) is extremely unintuitive, difficult to express and inefficient to compute. Hence, we propose an extension to the ANSI SQL join operator, referred to as "PARTITIONED OUTER JOIN", which allows for a succinct expression of densification along the dimensions of interest. We also present various algorithms to evaluate the new join operator efficiently and compare it with existing methods of doing the equivalent operation. We also define a new window function "LAST_VALUE (IGNORE NULLS)" which is very useful with partitioned outer join.
Abhinav Gupta 0003, Sankar Subramanian, Srikanth Bellamkonda, Tolga Bozkaya, Nathan Folkert, Andrew Witkowski
SIGMOD Conference7
2004 Where is Business Intelligence taking today's Database Systems?
William O'Connell, Andrew Witkowski, Ramesh Bhashyam, Surajit Chaudhuri
VLDB2
2003 Spreadsheets in RDBMS for OLAP
abstract
One of the critical deficiencies of SQL is lack of support for n-dimensional array-based computations which are frequent in OLAP environments. Relational OLAP (ROLAP) applications have to emulate them using joins, recently introduced SQL Window Functions [18] and complex and inefficient CASE expressions. The designated place in SQL for specifying calculations is the SELECT clause, which is extremely limiting and forces the user to generate queries using nested views, subqueries and complex joins. Furthermore, SQL-query optimizer is pre-occupied with determining efficient join orders and choosing optimal access methods and largely disregards optimization of complex numerical formulas. Execution methods concentrated on efficient computation of a cube [11], [16] rather than on random access structures for inter-row calculations. This has created a gap that has been filled by spreadsheets and specialized MOLAP engines, which are good at formulas for mathematical modeling but lack the formalism of the relational model, are difficult to manage, and exhibit scalability problems. This paper presents SQL extensions involving array based calculations for complex modeling. In addition, we present optimizations, access structures and execution models for processing them efficiently. 1
Andrew Witkowski, Srikanth Bellamkonda, Tolga Bozkaya, Gregory Dorman, Nathan Folkert, Abhinav Gupta 0003, Sankar Subramanian
SIGMOD Conference1
2003 Business Modelling Using SQL Spreadsheets
Andrew Witkowski, Srikanth Bellamkonda, Tolga Bozkaya, Nathan Folkert, Abhinav Gupta 0003, Sankar Subramanian
VLDB1
2001 Collaborative Analytical Processing - Dream or Reality? (Panel abstract)
William O'Connell, Andrew Witkowski, Goetz Graefe
VLDB2
1998 Materialized Views in Oracle
Randall G. Bello, Karl Dias, Alan Downing, James J. Feenan Jr., James L. Finnerty, William D. Norcott, Harry Sun, Andrew Witkowski, Mohamed Ziauddin
VLDB8
1993 NCR 3700 - The Next-Generation Industrial Database Computer
Andrew Witkowski, Felipe Cariño, Pekka Kostamaa
VLDB1
1986 Performance Evaluation of Multiversion with the Oracle Synchronization
abstract
In this paper we present a new analytical model for performance measurements of timestamp driven databases. The model is based on two-dimensional Poisson processes where one coordinate represents the real arrival time and the other the timestamp of an arriving messages. The notion of preemption is defined which serves as a model for synchronization. Preemption naturally implies such performance measures as response time and amount of abortion in the system. The concept of oracle is introduced which allows evaluation of a lower bound on the synchronization cost. Preemption and the oracle are then used to evaluate performance of the Multiversion synchronization. We present the distribution and the expectation of the synchronization cost. The analysis is then applied to a database with exponential communication delays (a) and the intensity of transaction l. It is shown that for Multiversion, this cost depends linearly on l/a and logarithmically on l.
Andrew Witkowski
SIGMETRICS1
1984 An Approach to Performance Analysis of Timestamp-driven Synchronization Mechanisms
abstract
In this paper we introduce a new analytical approach to modeling the performance of systems synchronized by timestamp mechanisms, including database systems. We define the virtual time - real time (T-V) plane, and an important kind of stochastic process that we call linear Poisson processes. We show how to calculate the rate of preemption (corresponding to the rate of abortion or rollback in concurrency control mechanisms) and the waiting time until last preemption (corresponding to commit time) for linear Poisson processes. Finally, we apply this theory, analyzing one example system synchronized by the Time Warp mechanism.
David Jefferson, Andrew Witkowski
PODC2