Cristian Bârca

dblp:181/5876 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
0since 2021 · last 2016
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Distributed and cloud data management · 44% Query processing and optimization · 44% Transaction processing and concurrency control · 13%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
parallel query processing
0.212016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Distributed and cloud data management › big data systems
SQL-on-Hadoop
0.212016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Transaction processing and concurrency control
update management
0.112016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Cloud and datacenter computing › resource management
workload management
0.112016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016

Methods — techniques the papers use, named apart from their topics

positional delta trees · 0.5HDFS replication policy · 0.5
YearPublicationVenuePosition
2016 VectorH: Taking SQL-on-Hadoop to the Next Level
abstract
Actian Vector in Hadoop (VectorH for short) is a new SQL-on-Hadoop system built on top of the fast Vectorwise analytical database system. VectorH achieves fault tolerance and storage scalability by relying on HDFS, and extends the state-of-the-art in SQL-on-Hadoop systems by instrumenting the HDFS replication policy to optimize read locality. VectorH integrates with YARN for workload management, achieving a high degree of elasticity. Even though HDFS is an append-only filesystem, and VectorH supports (update-averse) ordered tables, trickle updates are possible thanks to Positional Delta Trees (PDTs), a differential update structure that can be queried efficiently. We describe the changes made to single-server Vectorwise to turn it into a Hadoop-based MPP system, encompassing workload management, parallel query optimization and execution, HDFS storage, transaction processing and Spark integration. We evaluate VectorH against HAWQ, Impala, SparkSQL and Hive, showing orders of magnitude better performance.
Andrei Costea, Adrian Ionescu, Bogdan Raducanu, Michal Switakowski, Cristian Bârca, Juliusz Sompolski, Alicja Luszczak, Michal Szafranski, Giel de Nijs, Peter Boncz
SIGMOD Conference5