David E. Simmen

dblp:64/3571 · DBLP profile ↗
← Back
11ranked-venue papers
5as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 11 · 5 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
10 papers
Query processing and optimization · 52% Database system architecture and tuning · 23% Data integration and cleaning · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 10 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query optimization
0.222015
Accelerating Big Data analytics with Collaborative Planning in Teradata Aster 6 · ICDE 2015
Using EELs, a Practical Approach to Outerjoin and Antijoin Reordering · ICDE 2001
Graph data management › graph analytics
large-scale graph analytics
0.212014
Large-Scale Graph Analytics in Aster 6: Bringing Context to Big Data Discovery · Proc. VLDB Endow. 2014
Data integration and cleaning › web data integration
data mashup
0.222008
Damia: data mashups for intranet applications · SIGMOD Conference 2008
DAMIA - A Data Mashup Fabric for Intranet Applications · VLDB 2007
Query processing and optimization › adaptive query processing
adaptive query optimization
0.012004
Progressive Optimization in Action · VLDB 2004
Query processing and optimization
cardinality estimation
0.012004
Robust Query Processing through Progressive Optimization · SIGMOD Conference 2004
Query processing and optimization › adaptive query processing
progressive optimization
0.012004
Progressive Optimization in Action · VLDB 2004
Query processing and optimization › adaptive query processing
query re-optimization
0.012004
Robust Query Processing through Progressive Optimization · SIGMOD Conference 2004
Query processing and optimization › query optimization
join ordering
0.012001
Using EELs, a Practical Approach to Outerjoin and Antijoin Reordering · ICDE 2001
Query processing and optimization › query optimization › join ordering
outerjoin and antijoin reordering
0.012001
Using EELs, a Practical Approach to Outerjoin and Antijoin Reordering · ICDE 2001
Query processing and optimization › query execution
sort optimization
0.011996
Fundamental Techniques for Order Optimization · SIGMOD Conference 1996

Methods — techniques the papers use, named apart from their topics

mapreduce · 0.7query planning · 0.5graph processing · 0.5graph execution · 0.2information extraction · 0.2vertex-oriented API · 0.2bulk synchronous parallel execution · 0.2SQL-MR integration · 0.2sensitivity analysis · 0.0checkpoint operators · 0.0
YearPublicationVenuePosition
2016 SQL-SA for big data discovery polymorphic and parallelizable SQL user-defined scalar and aggregate infrastructure in Teradata Aster 6.20
abstract
There is increasing demand to integrate big data analytic systems using SQL. Given the vast ecosystem of SQL applications, enabling SQL capabilities allows big data platforms to expose their analytic potential to a wide variety of end users, accelerating discovery processes and providing significant business value. Most existing big data frameworks are based on one particular programming model such as MapReduce or Graph. However, data scientists are often forced to manually create adhoc data pipelines to connect various big data tools and platforms to serve their analytic needs. When the analytic tasks change, these data pipelines may be costly to modify and maintain. In this paper we present SQL-SA, a polymorphic and parallelizable SQL scalar and aggregate infrastructure in Aster 6.20. This infrastructure extends Aster 6's MapReduce and Graph capabilities to support polymorphic user-defined scalar and aggregate functions using flexible SQL syntax. The implementation enhances main Aster components including query syntax, API, planning and execution extensively. Integrating these new user-defined scalar and aggregate functions with Aster MapReduce and Graph functions, Aster 6.20 enables data scientists to integrate diverse programming models in a single SQL statement. The statement is automatically converted to an optimal data pipeline and executed in parallel. Using a real world business problem and data, Aster 6.20 demonstrates a significant performance advantage (25%+) over Hadoop Pig and Hive.
Robert M. Wehrmeister, James Shau, Abhirup Chakraborty, Daley Alex, Awny Al Omari, Feven Atnafu, Jeff Davis, Litao Deng, Deepak Jaiswal, Chittaranjan Keswani, Yafeng Lu, Tom Reyes, Kashif Siddiqui, David E. Simmen, Devendra Vidhani, Daniel Yu
ICDE16
2015 Accelerating Big Data analytics with Collaborative Planning in Teradata Aster 6
abstract
The volume, velocity, and variety of Big Data necessitate the development of new and innovative data processing software. A multitude of SQL implementations on distributed systems have emerged in recent years to enable large-scale data analysis. User-Defined Table operators (written in procedural languages) embedded in these SQL implementations are a powerful mechanism to succinctly express and perform analytic operations typical in Big Data discovery workloads. Table operators can be easily customized to implement different processing models such as map, reduce and graph execution. Despite an inherently parallel execution model, the performance and scalability of these table operators is greatly restricted as they appear as a black box to a typical SQL query optimizer. The optimizer is not able to infer even the basic properties of table operators, prohibiting the application of optimization rules and strategies. In this paper, we introduce an innovative concept of “Collaborative Planning”, which results in the removal of redundant operations and a more optimal rearrangement of query plan operators. The optimization of the query proceeds through a collaborative exchange between the planner and the table operator. Plan properties and context information of surrounding query plan operations are exchanged between the optimizer and the table operator. Knowing these properties also allows the author of the table operator to optimize its embedded logic. Our main contribution in this paper is the design and implementation of Collaborative Planning in the Teradata Aster 6 system. Using real-world workloads, we show that Collaborative Planning reduces query execution times as much as 90.0% in common use cases, resulting in a 24x speedup.
Aditi Pandit, Derrick Kondo, David E. Simmen, Anjali Norwood, Tongxin Bai
ICDE3
2014 Large-Scale Graph Analytics in Aster 6: Bringing Context to Big Data Discovery
abstract
Graph analytics is an important big data discovery technique. Applications include identifying influential employees for retention, detecting fraud in a complex interaction network, and determining product affinities by exploiting community buying patterns. Specialized platforms have emerged to satisfy the unique processing requirements of large-scale graph analytics; however, these platforms do not enable graph analytics to be combined with other analytics techniques, nor do they work well with the vast ecosystem of SQL-based business applications. Teradata Aster 6.0 adds support for large-scale graph analytics to its repertoire of analytics capabilities. The solution extends the multi-engine processing architecture with support for bulk synchronous parallel execution, and a specialized graph engine that enables iterative analysis of graph structures. Graph analytics functions written to the vertex-oriented API exposed by the graph engine can be invoked from the context of an SQL query and composed with existing SQL-MR functions, thereby enabling data scientists and business applications to express computations that combine large-scale graph analytics with techniques better suited to a different style of processing. The solution includes a suite of pre-built graph analytic functions adapted for parallel execution.
David E. Simmen, Karl Schnaitter, Jeff Davis, Sangeet Lohariwala, Ajay Mysore, Vinayak Shenoi, Mingfeng Tan
Proc. VLDB Endow.1
2009 Enabling enterprise mashups over unstructured text feeds with InfoSphere MashupHub and SystemT
abstract
Enterprise mashup scenarios often involve feeds derived from data created primarily for eye consumption, such as email, news, calendars, blogs, and web feeds. These data sources can test the capabilities of current data mashup products, as the attributes needed to perform join, aggregation, and other operations are often buried within unstructured feed text. Information extraction technology is a key enabler in such scenarios, using annotators to convert unstructured text into structured information that can facilitate mashup operations.
David E. Simmen, Frederick Reiss 0001, Yunyao Li 0001, Suresh Thalamati
SIGMOD Conference1
2008 Damia: data mashups for intranet applications
abstract
Increasingly large numbers of situational applications are being created by enterprise business users as a by-product of solving day-to-day problems. In efforts to address the demand for such applications, corporate IT is moving toward Web 2.0 architectures. In particular, the corporate intranet is evolving into a platform of readily accessible data and services where communities of business users can assemble and deploy situational applications. Damia is a web style data integration platform being developed to address the data problem presented by such applications, which often access and combine data from a variety of sources. Damia allows business users to quickly and easily create data mashups that combine data from desktop, web, and traditional IT sources into feeds that can be consumed by AJAX, and other types of web applications. This paper describes the key features and design of Damia's data integration engine, which has been packaged with Mashup Hub, an enterprise feed server currently available for download on IBM alphaWorks. Mashup Hub exposes Damia's data integration capabilities in the form of a service that allows users to create hosted data mashups.
David E. Simmen, Mehmet Altinel, Volker Markl, Sriram Padmanabhan
SIGMOD Conference1
2007 DAMIA - A Data Mashup Fabric for Intranet Applications
Mehmet Altinel, Paul Brown, Susan Cline, Rajesh Kartha, Eric Louie, Volker Markl, Louis Mau, Yip-Hing Ng, David E. Simmen
VLDB9
2004 Robust Query Processing through Progressive Optimization
abstract
Virtually every commercial query optimizer chooses the best plan for a query using a cost model that relies heavily on accurate cardinality estimation. Cardinality estimation errors can occur due to the use of inaccurate statistics, invalid assumptions about attribute independence, parameter markers, and so on. Cardinality estimation errors may cause the optimizer to choose a sub-optimal plan. We present an approach to query processing that is extremely robust because it is able to detect and recover from cardinality estimation errors. We call this approach "progressive query optimization" (POP). POP validates cardinality estimates against actual values as measured during query execution. If there is significant disagreement between estimated and actual values, execution might be stopped and re-optimization might occur. Oscillation between optimization and execution steps can occur any number of times. A re-optimization step can exploit both the actual cardinality and partial results, computed during a previous execution step. Checkpoint operators (CHECK) validate the optimizer's cardinality estimates against actual cardinalities. Each CHECK has a condition that indicates the cardinality bounds within which a plan is valid. We compute this validity range through a novel sensitivity analysis of query plan operators. If the CHECK condition is violated, CHECK triggers re-optimization. POP has been prototyped in a leading commercial DBMS. An experimental evaluation of POP using TPC-H queries illustrates the robustness POP adds to query processing, while incurring only negligible overhead. A case-study applying POP to a real-world database and workload shows the potential of POP, accelerating complex OLAP queries by almost two orders of magnitude.
Volker Markl, Vijayshankar Raman, David E. Simmen, Guy M. Lohman, Hamid Pirahesh
SIGMOD Conference3
2004 Progressive Optimization in Action
Vijayshankar Raman, Volker Markl, David E. Simmen, Guy M. Lohman, Hamid Pirahesh
VLDB3
2001 Using EELs, a Practical Approach to Outerjoin and Antijoin Reordering
abstract
Outerjoins and antijoins are two important classes of joins in database systems. Reordering outerjoins and antijoins with innerjoins is challenging because not all the join orders preserve the semantics of the original query. Previous work did not consider antijoins and was restricted to a limited class of queries. We consider using a conventional bottom-up optimizer to reorder different types of joins. We propose extending each join predicate's eligibility list, which contains all the tables referenced in the predicate. An extended eligibility list (EEL) includes all the tables needed by a predicate to preserve the semantics of the original query. We describe an algorithm that can set up the EELs properly in a bottom-up traversal of the original operator tree. A conventional join optimizer is then modified to check the EELs when generating sub-plans. Our approach handles antijoin and can resolve many practical issues. It is now being implemented in an upcoming release of IBM's Universal Database Server for Unix, Windows and OS/2.
Jun Rao, Bruce G. Lindsay 0001, Guy M. Lohman, Hamid Pirahesh, David E. Simmen
ICDE5
1996 Fundamental Techniques for Order Optimization
David E. Simmen, Eugene J. Shekita, Timothy Malkemus
EDBT1
1996 Fundamental Techniques for Order Optimization
abstract
Decision support applications are growing in popularity as more business data is kept on-line. Such applications typically include complex SQL queries that can test a query optimizer's ability to produce an efficient access plan. Many access plan strategies exploit the physical ordering of data provided by indexes or sorting. Sorting is an expensive operation, however. Therefore, it is imperative that sorting is optimized in some way or avoided all together. Toward that goal, this paper describes novel optimization techniques for pushing down sorts in joins, minimizing the number of sorting columns, and detecting when sorting can be avoided because of predicates, keys, or indexes. A set of fundamental operations is described that provide the foundation for implementing such techniques. The operations exploit data properties that arise from predicate application, uniqueness, and functional dependencies. These operations and techniques have been implemented in IBM's DB2/CS.
David E. Simmen, Eugene J. Shekita, Timothy Malkemus
SIGMOD Conference1