Ariel Cary

dblp:77/3857 · DBLP profile ↗
← Back
5ranked-venue papers
2as first author
0since 2021 · last 2014
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 5 · 2 first-authorArtificial intelligence and machine learning · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
2 papers
Query processing and optimization · 42% Database system architecture and tuning · 24% Data models and query languages · 21%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Database system architecture and tuning › database design
physical database design
0.212014
DBDesigner: A customizable physical design tool for Vertica Analytic Database · ICDE 2014
Query processing and optimization › materialization
late materialization
0.212013
Materialization strategies in the Vertica analytic database: Lessons learned · ICDE 2013
Data models and query languages › datalog › datalog query optimization
sideways information passing
0.212013
Materialization strategies in the Vertica analytic database: Lessons learned · ICDE 2013
Indexing and storage engines
column store
0.012013
Materialization strategies in the Vertica analytic database: Lessons learned · ICDE 2013

Methods — techniques the papers use, named apart from their topics

optimizer cost estimation · 0.2cost-benefit model · 0.2experimental comparison · 0.2
YearPublicationVenuePosition
2014 DBDesigner: A customizable physical design tool for Vertica Analytic Database
abstract
In this paper, we present Vertica's customizable physical design tool, called the DBDesigner (DBD), that produces designs optimized for various scenarios and applications. For a given workload and space budget, DBD automatically recommends a physical design that optimizes query performance, storage footprint, fault tolerance and recovery to meet different customer requirements. Vertica is a distributed, massively parallel columnar database that physically organizes data into projections. Projections are attribute subsets from one or more tables with tuples sorted by one or more attributes, that are replicated or segmented (distributed) on cluster nodes. The key challenges involved in projection design are picking appropriate column sets, sort orders, cluster data distributions and column encodings. To achieve the desired trade-off between query performance and storage footprint, DBD operates under three different design policies: (a) load-optimized, (b) query-optimized or (c) balanced. These policies indirectly control the number of projections proposed and queries optimized to achieve the desired balance. To cater to query workloads that evolve over time, DBD also operates in a comprehensive and incremental design mode. In addition, DBD lets users override specific features of projection design based on their intimate knowledge about the data and query workloads. We present the complete physical design algorithm, describing in detail how projection candidates are efficiently explored and evaluated using optimizer's cost and benefit model. Our experimental results show that DBD produces good physical designs that satisfy a variety of customer use cases.
Ramakrishna Varadarajan, Vivek Bharathan, Ariel Cary, Jaimin Dave, Sreenath Bodagala
ICDE3
2013 Materialization strategies in the Vertica analytic database: Lessons learned
abstract
Column store databases allow for various tuple reconstruction strategies (also called materialization strategies). Early materialization is easy to implement but generally performs worse than late materialization. Late materialization is more complex to implement, and usually performs much better than early materialization, although there are situations where it is worse. We identify these situations, which essentially revolve around joins where neither input fits in memory (also called spilling joins). Sideways information passing techniques provide a viable solution to get the best of both worlds. We demonstrate how early materialization combined with sideways information passing allows us to get the benefits of late materialization, without the bookkeeping complexity or worse performance for spilling joins. It also provides some other benefits to query processing in Vertica due to positive interaction with compression and sort orders of the data. In this paper, we report our experiences with late and early materialization, highlight their strengths and weaknesses, and present the details of our sideways information passing implementation. We show experimental results of comparing these materialization strategies, which highlight the significant performance improvements provided by our implementation of sideways information passing (up to 72% on some TPC-H queries).
Lakshmikant Shrinivas, Sreenath Bodagala, Ramakrishna Varadarajan, Ariel Cary, Vivek Bharathan, Chuck Bear
ICDE4
2011 SpSJoin: parallel spatial similarity joins
abstract
A spatial similarity join of two geospatial datasets finds pairs of records that are simultaneously similar on spatial and textual attributes. Such join is useful for a variety of applications, like data cleansing, record linkage, duplications detection and geocoding enhancement. Efficient techniques exist for the individual joins on either spatial or textual attributes. However, the combined problem has received much less research attention. This paper presents the SpSJoin (Spatial Similarity join) system to fill in this need. SpSJoin is a platform that merges geospatial and text processing techniques for efficiently performing spatial similarity joins. The platform leverages parallel computing with MapReduce to tackle scalability issues in joining large datasets. The efficiency of the proposed techniques are experimentally validated with a join case for improving the geolocation of entities in a real geospatial dataset with referential entities of another dataset.
Jaime Ballesteros, Ariel Cary, Naphtali Rishe
GIS2
2010 Efficient and Scalable Method for Processing Top-k Spatial Boolean Queries
Ariel Cary, Ouri Wolfson, Naphtali Rishe
SSDBM1
2009 Experiences on Processing Spatial Data with MapReduce
Ariel Cary, Zhengguo Sun, Vagelis Hristidis, Naphtali Rishe
SSDBM1