VLDB 2026 Research / reviewers in the wild / expert
Alan Gates
dblp:71/7259 · also Alan F. Gates
· DBLP profile ↗
3ranked-venue papers
1as first author
0since 2021 · last 2019
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 3 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
3 papers |
Query processing and optimization · 28% Database system architecture and tuning · 22% Transaction processing and concurrency control · 22% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Storage systems · 100% |
Topics — the 6 heaviest of 9, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Transaction processing and concurrency control
ACID transactions |
0.4 | 1 | 2019 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing · SIGMOD Conference 2019 |
Query processing and optimization
query optimization |
0.4 | 1 | 2019 | Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing · SIGMOD Conference 2019 |
Data integration and cleaning
data warehouse |
0.2 | 1 | 2014 | Major technical advancements in apache hive · SIGMOD Conference 2014 |
Storage systems › data management › database storage
columnar storage |
0.2 | 1 | 2014 | Major technical advancements in apache hive · SIGMOD Conference 2014 |
Storage systems › data representation
file format |
0.2 | 1 | 2014 | Major technical advancements in apache hive · SIGMOD Conference 2014 |
Query processing and optimization
query compilation |
0.1 | 1 | 2009 | Building a HighLevel Dataflow System on top of MapReduce: The Pig Experience · Proc. VLDB Endow. 2009 |
Methods — techniques the papers use, named apart from their topics
mapreduce · 0.4MPP · 0.4benchmarking · 0.4custom map/reduce functions · 0.1SQL-style constructs · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2019 | Apache Hive: From MapReduce to Enterprise-grade Big Data WarehousingabstractApache Hive is an open-source relational database system for analytic big-data workloads. In this paper we describe the key innovations on the journey from batch tool to fully fledged enterprise data warehousing system. We present a hybrid architecture that combines traditional MPP techniques with more recent big data and cloud concepts to achieve the scale and performance required by today's analytic applications. We explore the system by detailing enhancements along four main axis: Transactions, optimizer, runtime, and federation. We then provide experimental results to demonstrate the performance of the system for typical workloads and conclude with a look at the community roadmap. Jesús Camacho-Rodríguez, Ashutosh Chauhan, Alan Gates, Eugene Koifman, Owen O'Malley, Vineet Garg, Zoltan Haindrich, Sergey Shelukhin, Prasanth Jayachandran, Siddharth Seth, Deepak Jaiswal, Slim Bouguerra, Nishant Bangarwa, Sankar Hariappan, Anishek Agarwal, Jason Dere, Daniel Dai, Thejas Nair, Nita Dembla, Gopal Vijayaraghavan, Günther Hagleitner |
SIGMOD Conference | 3 |
| 2014 | Major technical advancements in apache hiveabstractApache Hive is a widely used data warehouse system for Apache Hadoop, and has been adopted by many organizations for various big data analytics applications. Closely working with many users and organizations, we have identified several shortcomings of Hive in its file formats, query planning, and query execution, which are key factors determining the performance of Hive. In order to make Hive continuously satisfy the requests and requirements of processing increasingly high volumes data in a scalable and efficient way, we have set two goals related to storage and runtime performance in our efforts on advancing Hive. First, we aim to maximize the effective storage capacity and to accelerate data accesses to the data warehouse by updating the existing file formats. Second, we aim to significantly improve cluster resource utilization and runtime performance of Hive by developing a highly optimized query planner and a highly efficient query execution engine. In this paper, we present a community-based effort on technical advancements in Hive. Our performance evaluation shows that these advancements provide significant improvements on storage efficiency and query execution performance. This paper also shows how academic research lays a foundation for Hive to improve its daily operations. Yin Huai, Ashutosh Chauhan, Alan Gates, Günther Hagleitner, Eric N. Hanson, Owen O'Malley, Jitendra Pandey, Yuan Yuan 0014, Rubao Lee, Xiaodong Zhang 0001 |
SIGMOD Conference | 3 |
| 2009 | Building a HighLevel Dataflow System on top of MapReduce: The Pig ExperienceabstractIncreasingly, organizations capture, transform and analyze enormous data sets. Prominent examples include internet companies and e-science. The Map-Reduce scalable dataflow paradigm has become popular for these applications. Its simple, explicit dataflow programming model is favored by some over the traditional high-level declarative approach: SQL. On the other hand, the extreme simplicity of Map-Reduce leads to much low-level hacking to deal with the many-step, branching dataflows that arise in practice. Moreover, users must repeatedly code standard operations such as join by hand. These practices waste time, introduce bugs, harm readability, and impede optimizations. Pig is a high-level dataflow system that aims at a sweet spot between SQL and Map-Reduce. Pig offers SQL-style high-level data manipulation constructs, which can be assembled in an explicit dataflow and interleaved with custom Map- and Reduce-style functions or executables. Pig programs are compiled into sequences of Map-Reduce jobs, and executed in the Hadoop Map-Reduce environment. Both Pig and Hadoop are open-source projects administered by the Apache Software Foundation. This paper describes the challenges we faced in developing Pig, and reports performance comparisons between Pig execution and raw Map-Reduce execution. Alan Gates, Olga Natkovich, Shubham Chopra, Pradeep Kamath, Shravan Narayanam, Christopher Olston, Benjamin C. Reed, Santhosh Srinivasan, Utkarsh Srivastava |
Proc. VLDB Endow. | 1 |