Vuk Ercegovac

dblp:58/4208 · DBLP profile ↗
← Back
16ranked-venue papers
2as first author
2since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Query processing and optimization · 65% Distributed and cloud data management · 13% Data models and query languages · 8%
Computer architecture, parallel and distributed computing, and storage systems
5 papers
Cloud and datacenter computing · 54% High-performance computing · 30% Parallel and multicore computing · 9%

Topics — the 19 heaviest of 29, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
adaptive query processing
0.812024
Adaptive and Robust Query Execution for Lakehouses At Scale · Proc. VLDB Endow. 2024
Distributed and cloud data management › mapreduce
mapreduce query processing
0.212015
Groupwise analytics via adaptive MapReduce · ICDE 2015
Query processing and optimization › adaptive query processing
dynamic query optimization
0.212014
Dynamically optimizing queries over large scale data platforms · SIGMOD Conference 2014
Query processing and optimization
query optimization
0.212014
Dynamically optimizing queries over large scale data platforms · SIGMOD Conference 2014
Cloud and datacenter computing
cluster resource management and scheduling
0.212013
PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics · Proc. VLDB Endow. 2013
Query processing and optimization
partial evaluation
0.112012
Declarative error management for robust data-intensive applications · SIGMOD Conference 2012
Query processing and optimization
query execution
0.112012
Declarative error management for robust data-intensive applications · SIGMOD Conference 2012
Query processing and optimization
join processing
0.112010
A comparison of join algorithms for log processing in MaPreduce · SIGMOD Conference 2010
Query processing and optimization › join processing › distributed join
mapreduce join
0.112010
A comparison of join algorithms for log processing in MaPreduce · SIGMOD Conference 2010
Data models and query languages
uncertain data management
0.112009
E = MC3: managing uncertain enterprise data in a cluster-computing environment · SIGMOD Conference 2009
High-performance computing › data-intensive computing
large-scale data processing
0.112014
Dynamically optimizing queries over large scale data platforms · SIGMOD Conference 2014
Performance modeling and evaluation
benchmarking
0.112005
The TEXTURE Benchmark: Measuring Performance of Text Queries on a Relational DBMS · VLDB 2005
Privacy and data protection › statistical database privacy
database privacy
0.012004
Limiting Disclosure in Hippocratic Databases · VLDB 2004
Privacy and data protection › privacy engineering
hippocratic databases
0.012004
Limiting Disclosure in Hippocratic Databases · VLDB 2004
Information retrieval
retrieval models
0.012003
The QUIQ Engine: A Hybrid IR DB System · ICDE 2003
Data mining
log analysis
0.012010
A comparison of join algorithms for log processing in MaPreduce · SIGMOD Conference 2010
Distributed and cloud data management
mapreduce
0.012010
A comparison of join algorithms for log processing in MaPreduce · SIGMOD Conference 2010
Database system architecture and tuning › database security
privacy-preserving databases
0.012004
Limiting Disclosure in Hippocratic Databases · VLDB 2004
Interaction techniques and input
direct manipulation
0.011998
DataSplash · SIGMOD Conference 1998

Methods — techniques the papers use, named apart from their topics

runtime statistics collection · 1.5incremental updating · 0.4data structure sharing · 0.4adaptive mapreduce · 0.4user-defined functions · 0.4cost-based optimization · 0.4monte carlo · 0.2mapreduce · 0.2sample runs · 0.2convergence trend modeling · 0.2declarative error handling · 0.1performance evaluation · 0.0
YearPublicationVenuePosition
2024 Adaptive and Robust Query Execution for Lakehouses At Scale
abstract
Many organizations have embraced the "Lakehouse" data management paradigm, which involves constructing structured data warehouses on top of open, unstructured data lakes. This approach stands in stark contrast to traditional, closed, relational databases and introduces challenges for performance and stability of distributed query processors. Firstly, in large-scale, open Lakehouses with uncurated data, high ingestion rates, external tables, or deeply nested schemas, it is often costly or wasteful to maintain perfect and up-to-date table and column statistics. Secondly, inherently imperfect cardinality estimates with conjunctive predicates, joins and user-defined functions can lead to bad query plans. Thirdly, for the sheer magnitude of data involved, strictly relying on static query plan decisions can result in performance and stability issues such as excessive data movement, substantial disk spillage, or high memory pressure. To address these challenges, this paper presents our design, implementation, evaluation and practice of the Adaptive Query Execution (AQE) framework, which exploits natural execution pipeline breakers in query plans to collect accurate statistics and re-optimize them at runtime for both performance and robustness. In the TPC-DS benchmark, the technique demonstrates up to 25× per query speedup. At Databricks, AQE has been successfully deployed in production for multiple years. It powers billions of queries and ETL jobs to process exabytes of data per day, through key enterprise products such as Databricks Runtime, Databricks SQL, and Delta Live Tables.
Maryann Xue, Yingyi Bu, Abhishek Somani, Wenchen Fan, Steven Chen, Herman Van Hövell, Bart Samwel, Mostafa Mokhtar, Rk Korlapati, Andy Lam, Yunxiao Ma, Vuk Ercegovac, Jiexing Li, Alexander Behm, Yuanjian Li, Xiao Li 0087, Sriram Krishnamurthy, Amit Shukla 0001, Michalis Petropoulos, Sameer Paranjpye, Reynold Xin, Matei Zaharia
Proc. VLDB Endow.13
2023 Making Data Engineering Declarative
Michael Armbrust, Ali Ghodsi 0002, Reynold Xin, Vuk Ercegovac, Sourav Chatterji, Eun-Gyu Kim, Paul Lappas, Yannis Papakonstantinou, Yingyi Bu, Yijia Cui, Rahul Govind, Aakash Japi, Kiavash Kianfar, Jon Mio, Mukul Murthy, Supun Nakandala, Yannis Sismanis, Justin Tang, Joseph Torres
CIDR4
2015 Groupwise analytics via adaptive MapReduce
abstract
Shared-nothing systems such as Hadoop vastly simplify parallel programming when processing disk-resident data whose size exceeds aggregate cluster memory. Such systems incur a significant performance penalty, however, on the important class of “groupwise set-valued analytics” (GSVA) queries in which the data is dynamically partitioned into groups and then a set-valued synopsis is computed for some or all of the groups. Key examples of synopses include top-k sets, bottom-k sets, and uniform random samples. Applications of GSVA queries include micro-marketing, root-cause analysis for problem diagnosis, and fraud detection. A naive approach to executing GSVA queries first reshuffles all of the data so that all records in a group are at the same node and then computes the synopsis for the group. This approach can be extremely inefficient when, as is typical, only a very small fraction of the records in each group actually contribute to the final groupwise synopsis, so that most of the shuffling effort is wasted. We show how to significantly speed up GSVA queries by slightly modifying the shared-nothing environment to allow tasks to occasionally access a small, common data structure; we focus on the Hadoop setting and use the “Adaptive MapReduce” infrastructure of Vernica et al. to implement the data structure. Our approach retains most of the advantages of a system such as Hadoop while significantly improving GSVA query performance, and also allows for incremental updating of query results. Experiments show speedups of up to 5x. Importantly, our new technique can potentially be applied to other shared-nothing systems with disk-resident data.
Liping Peng, Vuk Ercegovac, Kai Zeng 0002, Peter J. Haas, Andrey Balmin, Yannis Sismanis
ICDE2
2014 Dynamically optimizing queries over large scale data platforms
abstract
Enterprises are adapting large-scale data processing platforms, such as Hadoop, to gain actionable insights from their "big data". Query optimization is still an open challenge in this environment due to the volume and heterogeneity of data, comprising both structured and un/semi-structured datasets. Moreover, it has become common practice to push business logic close to the data via user-defined functions (UDFs), which are usually opaque to the optimizer, further complicating cost-based optimization. As a result, classical relational query optimization techniques do not fit well in this setting, while at the same time, suboptimal query plans can be disastrous with large datasets. In this paper, we propose new techniques that take into account UDFs and correlations between relations for optimizing queries running on large scale clusters. We introduce "pilot runs", which execute part of the query over a sample of the data to estimate selectivities, and employ a cost-based optimizer that uses these selectivities to choose an initial query plan. Then, we follow a dynamic optimization approach, in which plans evolve as parts of the queries get executed. Our experimental results show that our techniques produce plans that are at least as good as, and up to 2x (4x) better for Jaql (Hive) than, the best hand-written left-deep query plans.
Konstantinos Karanasos, Andrey Balmin, Marcel Kutsch, Fatma Özcan 0001, Vuk Ercegovac, Chunyang Xia, Jesse Jackson
SIGMOD Conference5
2013 PREDIcT: Towards Predicting the Runtime of Large Scale Iterative Analytics
abstract
Machine learning algorithms are widely used today for analytical tasks such as data cleaning, data categorization, or data filtering. At the same time, the rise of social media motivates recent uptake in large scale graph processing. Both categories of algorithms are dominated by iterative subtasks, i.e., processing steps which are executed repetitively until a convergence condition is met. Optimizing cluster resource allocations among multiple workloads of iterative algorithms motivates the need for estimating their runtime, which in turn requires: i) predicting the number of iterations, and ii) predicting the processing time of each iteration. As both parameters depend on the characteristics of the dataset and on the convergence function, estimating their values before execution is difficult. This paper proposes PREDIcT, an experimental methodology for predicting the runtime of iterative algorithms. PREDIcT uses sample runs for capturing the algorithm's convergence trend and per-iteration key input features that are well correlated with the actual processing requirements of the complete input dataset. Using this combination of characteristics we predict the runtime of iterative algorithms, including algorithms with very different runtime patterns among subsequent iterations. Our experimental evaluation of multiple algorithms on scale-free graphs shows a relative prediction error of 10%-30% for predicting runtime, including algorithms with up to 100× runtime variability among consecutive iterations.
Adrian Daniel Popescu, Andrey Balmin, Vuk Ercegovac, Anastasia Ailamaki
Proc. VLDB Endow.3
2012 Adaptive MapReduce using situation-aware mappers
abstract
We propose new adaptive runtime techniques for MapReduce that improve performance and simplify job tuning. We implement these techniques by breaking a key assumption of MapReduce that mappers run in isolation. Instead, our mappers communicate through a distributed meta-data store and are aware of the global state of the job. However, we still preserve the fault-tolerance, scalability, and programming API of MapReduce. We utilize these "situation-aware mappers" to develop a set of techniques that make MapReduce more dynamic: (a) Adaptive Mappers dynamically take multiple data partitions (splits) to amortize mapper start-up costs; (b) Adaptive Combiners improve local aggregation by maintaining a cache of partial aggregates for the frequent keys; (c) Adaptive Sampling and Partitioning sample the mapper outputs and use the obtained statistics to produce balanced partitions for the reducers. Our experimental evaluation shows that adaptive techniques provide up to 3x performance improvement, in some cases, and dramatically improve performance stability across the board.
Rares Vernica, Andrey Balmin, Kevin S. Beyer, Vuk Ercegovac
EDBT4
2012 Declarative error management for robust data-intensive applications
abstract
We present an approach to declaratively manage run-time errors in data-intensive applications. When large volumes of raw data meet complex third-party libraries, deterministic run-time errors become likely, and existing query processors typically stop without returning a result when a run-time error occurs. The ability to degrade gracefully in the presence of run-time errors, and partially execute jobs, is typically limited to specific operators such as bulkloading.
Carl-Christian Kanne, Vuk Ercegovac
SIGMOD Conference2
2011 Jaql: A Scripting Language for Large Scale Semistructured Data Analysis
Kevin S. Beyer, Vuk Ercegovac, Rainer Gemulla, Andrey Balmin, Mohamed Y. Eltabakh, Carl-Christian Kanne, Fatma Özcan 0001, Eugene J. Shekita
Proc. VLDB Endow.2
2010 A comparison of join algorithms for log processing in MaPreduce
abstract
The MapReduce framework is increasingly being used to analyze large volumes of data. One important type of data analysis done with MapReduce is log processing, in which a click-stream or an event log is filtered, aggregated, or mined for patterns. As part of this analysis, the log often needs to be joined with reference data such as information about users. Although there have been many studies examining join algorithms in parallel and distributed DBMSs, the MapReduce framework is cumbersome for joins. MapReduce programmers often use simple but inefficient algorithms to perform joins. In this paper, we describe crucial implementation details of a number of well-known join strategies in MapReduce, and present a comprehensive experimental comparison of these join techniques on a 100-node Hadoop cluster. Our results provide insights that are unique to the MapReduce platform and offer guidance on when to use a particular join algorithm on this platform.
Spyros Blanas, Jignesh M. Patel, Vuk Ercegovac, Jun Rao, Eugene J. Shekita, Yuanyuan Tian 0001
SIGMOD Conference3
2009 E = MC3: managing uncertain enterprise data in a cluster-computing environment
abstract
Modern enterprises must manage uncertain data for purposes of risk assessment and decisionmaking under uncertainty. The Monte Carlo approach embodied in the MCDB system of Jampani et al. is well suited for such a task. MCDB can support industrial strength business-intelligence queries over uncertain warehouse data. Moreover, MCDB's extensible approach to specifying uncertainty can also capture complex stochastic prediction models, allowing sophisticated ``what-if'' analyses within the DBMS. The MCDB computations can be highly CPU intensive, but offer the potential for massive parallelization. To realize this potential, we provide a new system, called MC3 (Monte Carlo Computation on a Cluster), that extends the MCDB approach to the map-reduce processing framework. MC3 can exploit the robustness and scalability of map-reduce, and can handle data stored in non-relational formats. We show how MCDB query plans over ``tuple bundles'' can be translated to sequences of map-reduce operations over nested data, and describe different parallelization schemes. We also provide and analyze several novel distributed algorithms for adding pseudorandom number seeds to tuple bundles. These algorithms ensure statistical correctness of the Monte-Carlo computations while minimizing the seed length. Our experiments show that MC3 can scale well for a variety of workloads.
Kevin S. Beyer, Vuk Ercegovac, Peter J. Haas, Eugene J. Shekita
SIGMOD Conference3
2008 Supporting sub-document updates and queries in an inverted index
abstract
Inverted indexes have become the standard indexing method for supporting search queries in a variety of content-based applications. Examples of such applications include enterprise document management, e-mail, web search, and social networks. One shortcoming in current inverted index designs is that they support only document-level updates, forcing a full document to be reindexed even if just part of it changes. This paper describes a new inverted index design that enables applications to break a document into semantically meaningful sub-documents or "sections". Each section of a document can be updated separately, but search queries can still work seamlessly across sections. Our index design is motivated by applications where there is metadata associated with each document that tends to be smaller and more frequently updated than the document's content, but at the same time, it is desireable to search the metadata and content with the same index structure. A novel self-optimizing query execution algorithm is described to efficiently join the sections of a document in the inverted index. Experimental results on TREC and patent data are provided, showing that sections can dramatically improve overall system throughput on a mixed workload of updates and queries.
Vuk Ercegovac, Vanja Josifovski, Maurício R. Mediano, Eugene J. Shekita
CIKM1
2005 The TEXTURE Benchmark: Measuring Performance of Text Queries on a Relational DBMS
Vuk Ercegovac, David J. DeWitt, Raghu Ramakrishnan 0001
VLDB1
2004 Mass Collaboration: A Case Study
Raghu Ramakrishnan 0001, Andrew Baptist, Vuk Ercegovac, Matt Hanselman, Navin Kabra, Amit Marathe, Uri Shaft
IDEAS3
2004 Limiting Disclosure in Hippocratic Databases
Kristen LeFevre, Rakesh Agrawal 0001, Vuk Ercegovac, Raghu Ramakrishnan 0001, Yirong Xu, David J. DeWitt
VLDB3
2003 The QUIQ Engine: A Hybrid IR DB System
abstract
For applications that involve rapidly changing textual data and also require traditional DBMS capabilities, current systems are unsatisfactory. We describe a hybrid IR-DB system that serves as the basis for the QUIQ-Connect product, a collaborative customer support application. We present a novel query paradigm and system architecture, along with performance results.
Navin Kabra, Raghu Ramakrishnan 0001, Vuk Ercegovac
ICDE3
1998 DataSplash
abstract
Database visualization is an area of growing importance as database systems become larger and more accessible. DataSplash is an easy-to-use, integrated environment for navigating, creating, and querying visual representations of data. We will demonstrate the three main components which make up the DataSplash environment: a navigation system, a direct-manipulation interface for creating and modifying visualizations, and a direct-manipulation visual query system.
Christopher Olston, Allison Woodruff, Alex Aiken, Michael Chu, Vuk Ercegovac, Mark Lin, Mybrid Spalding, Michael Stonebraker
SIGMOD Conference5