VLDB 2026 Research / reviewers in the wild / expert
Burak Yavuz
dblp:163/2321
· DBLP profile ↗
4ranked-venue papers
0as first author
0since 2021 · last 2020
0000-0002-5262-0346ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3Artificial intelligence and machine learning · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Storage systems · 60% High-performance computing · 23% Cloud and datacenter computing · 13% | |
| Databases, data mining, and information retrieval
2 papers |
Database system architecture and tuning · 36% Query processing and optimization · 36% Data stream processing · 28% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 12 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › object storage
cloud object store |
0.4 | 1 | 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020 |
Storage systems
key-value storage |
0.4 | 1 | 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020 |
Storage systems
storage reliability |
0.4 | 1 | 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object Stores · Proc. VLDB Endow. 2020 |
Query processing and optimization › incremental computation
incremental query processing |
0.3 | 1 | 2018 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark · SIGMOD Conference 2018 |
Machine learning › Efficient and distributed learning
machine learning libraries |
0.2 | 1 | 2016 | MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016 |
High-performance computing › numerical linear algebra
distributed matrix computation |
0.2 | 1 | 2016 | Matrix Computations and Optimization in Apache Spark · KDD 2016 |
High-performance computing › numerical linear algebra
singular value decomposition |
0.2 | 1 | 2016 | Matrix Computations and Optimization in Apache Spark · KDD 2016 |
Mathematical optimization › continuous optimization
convex optimization |
0.2 | 1 | 2016 | Matrix Computations and Optimization in Apache Spark · KDD 2016 |
Query processing and optimization
query compilation |
0.1 | 1 | 2018 | Structured Streaming: A Declarative API for Real-Time Applications in Apache Spark · SIGMOD Conference 2018 |
Cloud and datacenter computing
cluster computing framework |
0.1 | 1 | 2016 | Matrix Computations and Optimization in Apache Spark · KDD 2016 |
Cloud and datacenter computing
cluster resource management and scheduling |
0.1 | 1 | 2016 | MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016 |
Distributed systems
distributed data processing |
0.1 | 1 | 2016 | MLlib: Machine Learning in Apache Spark · J. Mach. Learn. Res. 2016 |
Methods — techniques the papers use, named apart from their topics
linear algebra primitives · 0.5distributed optimization · 0.5distributed matrix operations · 0.5convex optimization · 0.5incremental view maintenance · 0.3code generation · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2020 | Delta Lake: High-Performance ACID Table Storage over Cloud Object StoresabstractCloud object stores such as Amazon S3 are some of the largest and most cost-effective storage systems on the planet, making them an attractive target to store large data warehouses and data lakes. Unfortunately, their implementation as key-value stores makes it difficult to achieve ACID transactions and high performance: metadata operations such as listing objects are expensive, and consistency guarantees are limited. In this paper, we present Delta Lake, an open source ACID table storage layer over cloud object stores initially developed at Databricks. Delta Lake uses a transaction log that is compacted into Apache Parquet format to provide ACID properties, time travel, and significantly faster metadata operations for large tabular datasets (e.g., the ability to quickly search billions of table partitions for those relevant to a query). It also leverages this design to provide high-level features such as automatic data layout optimization, upserts, caching, and audit logs. Delta Lake tables can be accessed from Apache Spark, Hive, Presto, Redshift and other systems. Delta Lake is deployed at thousands of Databricks customers that process exabytes of data per day, with the largest instances managing exabyte-scale datasets and billions of objects. Michael Armbrust, Tathagata Das, Sameer Paranjpye, Reynold Xin, Shixiong Zhu, Ali Ghodsi 0002, Burak Yavuz, Mukul Murthy, Joseph Torres, Liwen Sun, Peter Boncz, Mostafa Mokhtar, Herman Van Hövell, Adrian Ionescu, Alicja Luszczak, Michal Switakowski, Takuya Ueshin, Xiao Li 0087, Michal Szafranski, Pieter Senster, Matei Zaharia |
Proc. VLDB Endow. | 7 |
| 2018 | Structured Streaming: A Declarative API for Real-Time Applications in Apache SparkabstractWith the ubiquity of real-time data, organizations need streaming systems that are scalable, easy to use, and easy to integrate into business applications. Structured Streaming is a new high-level streaming API in Apache Spark based on our experience with Spark Streaming. Structured Streaming differs from other recent streaming APIs, such as Google Dataflow, in two main ways. First, it is a purely declarative API based on automatically incrementalizing a static relational query (expressed using SQL or DataFrames), in contrast to APIs that ask the user to build a DAG of physical operators. Second, Structured Streaming aims to support end-to-end real-time applications that integrate streaming with batch and interactive analysis. We found that this integration was often a key challenge in practice. Structured Streaming achieves high performance via Spark SQL's code generation engine and can outperform Apache Flink by up to 2x and Apache Kafka Streams by 90x. It also offers rich operational features such as rollbacks, code updates, and mixed streaming/batch execution. We describe the system's design and use cases from several hundred production deployments on Databricks, the largest of which process over 1 PB of data per month. Michael Armbrust, Tathagata Das, Joseph Torres, Burak Yavuz, Shixiong Zhu, Reynold Xin, Ali Ghodsi 0002, Ion Stoica, Matei Zaharia |
SIGMOD Conference | 4 |
| 2016 | Matrix Computations and Optimization in Apache SparkabstractWe describe matrix computations available in the cluster programming framework, Apache Spark. Out of the box, Spark provides abstractions and implementations for distributed matrices and optimization routines using these matrices. When translating single-node algorithms to run on a distributed cluster, we observe that often a simple idea is enough: separating matrix operations from vector operations and shipping the matrix operations to be ran on the cluster, while keeping vector operations local to the driver. In the case of the Singular Value Decomposition, by taking this idea to an extreme, we are able to exploit the computational power of a cluster, while running code written decades ago for a single core. Another example is our Spark port of the popular TFOCS optimization package, originally built for MATLAB, which allows for solving Linear programs as well as a variety of other convex programs. We conclude with a comprehensive set of benchmarks for hardware accelerated matrix computations from the JVM, which is interesting in its own right, as many cluster programming frameworks use the JVM. The contributions described in this paper are already merged into Apache Spark and available on Spark installations by default, and commercially supported by a slew of companies which provide further services. Reza Bosagh Zadeh, Alexander Ulanov, Burak Yavuz, Li Pu, Shivaram Venkataraman, Evan Randall Sparks, Aaron Staple, Matei Zaharia |
KDD | 4 |
| 2016 | MLlib: Machine Learning in Apache SparkabstractApache Spark is a popular open-source platform for large-scale data processing that is well-suited for iterative machine learning tasks. In this paper we present MLlib, Spark's open- source distributed machine learning library. MLlib provides efficient functionality for a wide range of learning settings and includes several underlying statistical, optimization, and linear algebra primitives. Shipped with Spark, MLlib supports several languages and provides a high-level API that leverages Spark's rich ecosystem to simplify the development of end-to-end machine learning pipelines. MLlib has experienced a rapid growth due to its vibrant open-source community of over 140 contributors, and includes extensive documentation to support further growth and to let users quickly get up to speed. Joseph K. Bradley, Burak Yavuz, Evan Randall Sparks, Shivaram Venkataraman, Davies Liu, Jeremy Freeman, D. B. Tsai, Manish Amde, Sean Owen, Doris Xin, Reynold Xin, Michael J. Franklin, Reza Bosagh Zadeh, Matei Zaharia, Ameet Talwalkar |
J. Mach. Learn. Res. | 3 |