Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Adrian Ionescu

dblp:141/7572 · DBLP profile ↗
← Back
3ranked-venue papers
1as first author
0since 2021 · last 2020
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 2Theory of computation · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 82% Cloud and datacenter computing · 18%
Databases, data mining, and information retrieval
2 papers
Database system architecture and tuning · 43% Distributed and cloud data management · 25% Query processing and optimization · 25%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Storage systems › object storage
cloud object store
0.412020
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020
Storage systems
key-value storage
0.412020
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020
Storage systems
storage reliability
0.412020
Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020
Query processing and optimization
parallel query processing
0.212016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Distributed and cloud data management › big data systems
SQL-on-Hadoop
0.212016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Transaction processing and concurrency control
update management
0.112016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Cloud and datacenter computing › cluster resource management and scheduling
cluster resource management
0.112016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016
Cloud and datacenter computing › resource management
workload management
0.112016
VectorH: Taking SQL-on-Hadoop to the Next Level · SIGMOD Conference 2016

Methods — techniques the papers use, named apart from their topics

positional delta trees · 0.5HDFS replication policy · 0.5
YearPublicationVenuePosition
2020 Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores
abstract
Cloud object stores such as Amazon S3 are some of the largest and most cost-effective storage systems on the planet, making them an attractive target to store large data warehouses and data lakes. Unfortunately, their implementation as key-value stores makes it difficult to achieve ACID transactions and high performance: metadata operations such as listing objects are expensive, and consistency guarantees are limited. In this paper, we present Delta Lake, an open source ACID table storage layer over cloud object stores initially developed at Databricks. Delta Lake uses a transaction log that is compacted into Apache Parquet format to provide ACID properties, time travel, and significantly faster metadata operations for large tabular datasets (e.g., the ability to quickly search billions of table partitions for those relevant to a query). It also leverages this design to provide high-level features such as automatic data layout optimization, upserts, caching, and audit logs. Delta Lake tables can be accessed from Apache Spark, Hive, Presto, Redshift and other systems. Delta Lake is deployed at thousands of Databricks customers that process exabytes of data per day, with the largest instances managing exabyte-scale datasets and billions of objects.
Michael Armbrust, Tathagata Das, Sameer Paranjpye, Reynold Xin, Shixiong Zhu, Ali Ghodsi 0002, Burak Yavuz, Mukul Murthy, Joseph Torres, Liwen Sun, Peter Boncz, Mostafa Mokhtar, Herman Van Hövell, Adrian Ionescu, Alicja Luszczak, Michal Switakowski, Takuya Ueshin, Xiao Li 0087, Michal Szafranski, Pieter Senster, Matei Zaharia
Proc. VLDB Endow.14
2016 VectorH: Taking SQL-on-Hadoop to the Next Level
abstract
Actian Vector in Hadoop (VectorH for short) is a new SQL-on-Hadoop system built on top of the fast Vectorwise analytical database system. VectorH achieves fault tolerance and storage scalability by relying on HDFS, and extends the state-of-the-art in SQL-on-Hadoop systems by instrumenting the HDFS replication policy to optimize read locality. VectorH integrates with YARN for workload management, achieving a high degree of elasticity. Even though HDFS is an append-only filesystem, and VectorH supports (update-averse) ordered tables, trickle updates are possible thanks to Positional Delta Trees (PDTs), a differential update structure that can be queried efficiently. We describe the changes made to single-server Vectorwise to turn it into a Hadoop-based MPP system, encompassing workload management, parallel query optimization and execution, HDFS storage, transaction processing and Spark integration. We evaluate VectorH against HAWQ, Impala, SparkSQL and Hive, showing orders of magnitude better performance.
Andrei Costea, Adrian Ionescu, Bogdan Raducanu, Michal Switakowski, Cristian Bârca, Juliusz Sompolski, Alicja Luszczak, Michal Szafranski, Giel de Nijs, Peter Boncz
SIGMOD Conference2
2014 On the role of complementation in implicit language equations and relations
Adrian Ionescu, Ernst L. Leiss
J. Comput. Syst. Sci.1