Owen O'Malley

dblp:137/6735 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
0since 2021 · last 2019
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 3Systems, architecture and hardware · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
3 papers
Database system architecture and tuning · 25% Query processing and optimization · 25% Transaction processing and concurrency control · 25%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Storage systems · 62% Performance modeling and evaluation · 38%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Transaction processing and concurrency control
ACID transactions
0.412019
Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing · SIGMOD Conference 2019
Query processing and optimization
query optimization
0.412019
Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing · SIGMOD Conference 2019
Data integration and cleaning
data warehouse
0.212014
Major technical advancements in apache hive · SIGMOD Conference 2014
Storage systems › data management › database storage
columnar storage
0.212014
Major technical advancements in apache hive · SIGMOD Conference 2014
Storage systems › data representation
file format
0.212014
Major technical advancements in apache hive · SIGMOD Conference 2014
Performance modeling and evaluation
benchmarking
0.212013
Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters · Proc. VLDB Endow. 2013
Storage systems
data placement
0.212013
Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters · Proc. VLDB Endow. 2013
Performance modeling and evaluation › benchmarking
i/o benchmarking
0.212013
Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters · Proc. VLDB Endow. 2013

Methods — techniques the papers use, named apart from their topics

mapreduce · 0.4MPP · 0.4benchmarking · 0.4workload characterization · 0.3benchmarking tool · 0.3
YearPublicationVenuePosition
2019 Apache Hive: From MapReduce to Enterprise-grade Big Data Warehousing
abstract
Apache Hive is an open-source relational database system for analytic big-data workloads. In this paper we describe the key innovations on the journey from batch tool to fully fledged enterprise data warehousing system. We present a hybrid architecture that combines traditional MPP techniques with more recent big data and cloud concepts to achieve the scale and performance required by today's analytic applications. We explore the system by detailing enhancements along four main axis: Transactions, optimizer, runtime, and federation. We then provide experimental results to demonstrate the performance of the system for typical workloads and conclude with a look at the community roadmap.
Jesús Camacho-Rodríguez, Ashutosh Chauhan, Alan Gates, Eugene Koifman, Owen O'Malley, Vineet Garg, Zoltan Haindrich, Sergey Shelukhin, Prasanth Jayachandran, Siddharth Seth, Deepak Jaiswal, Slim Bouguerra, Nishant Bangarwa, Sankar Hariappan, Anishek Agarwal, Jason Dere, Daniel Dai, Thejas Nair, Nita Dembla, Gopal Vijayaraghavan, Günther Hagleitner
SIGMOD Conference5
2014 Major technical advancements in apache hive
abstract
Apache Hive is a widely used data warehouse system for Apache Hadoop, and has been adopted by many organizations for various big data analytics applications. Closely working with many users and organizations, we have identified several shortcomings of Hive in its file formats, query planning, and query execution, which are key factors determining the performance of Hive. In order to make Hive continuously satisfy the requests and requirements of processing increasingly high volumes data in a scalable and efficient way, we have set two goals related to storage and runtime performance in our efforts on advancing Hive. First, we aim to maximize the effective storage capacity and to accelerate data accesses to the data warehouse by updating the existing file formats. Second, we aim to significantly improve cluster resource utilization and runtime performance of Hive by developing a highly optimized query planner and a highly efficient query execution engine. In this paper, we present a community-based effort on technical advancements in Hive. Our performance evaluation shows that these advancements provide significant improvements on storage efficiency and query execution performance. This paper also shows how academic research lays a foundation for Hive to improve its daily operations.
Yin Huai, Ashutosh Chauhan, Alan Gates, Günther Hagleitner, Eric N. Hanson, Owen O'Malley, Jitendra Pandey, Yuan Yuan 0014, Rubao Lee, Xiaodong Zhang 0001
SIGMOD Conference6
2013 Apache Hadoop YARN: yet another resource negotiator
abstract
The initial design of Apache Hadoop [1] was tightly focused on running massive, MapReduce jobs to process a web crawl. For increasingly diverse companies, Hadoop has become the data and computational agorá---the de facto place where data and computational resources are shared and accessed. This broad adoption and ubiquitous usage has stretched the initial design well beyond its intended target, exposing two key shortcomings: 1) tight coupling of a specific programming model with the resource management infrastructure, forcing developers to abuse the MapReduce programming model, and 2) centralized handling of jobs' control flow, which resulted in endless scalability concerns for the scheduler.
Vinod Kumar Vavilapalli, Arun C. Murthy, Chris Douglas, Sharad Agarwal, Mahadev Konar, Robert Evans, Thomas Graves, Jason Lowe, Hitesh Shah, Siddharth Seth, Bikas Saha, Carlo Curino, Owen O'Malley, Sanjay Radia, Benjamin C. Reed, Eric Baldeschwieler
SoCC13
2013 Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters
abstract
A table placement method is a critical component in big data analytics on distributed systems. It determines the way how data values in a two-dimensional table are organized and stored in the underlying cluster. Based on Hadoop computing environments, several table placement methods have been proposed and implemented. However, a comprehensive and systematic study to understand, to compare, and to evaluate different table placement methods has not been done. Thus, it is highly desirable to gain important insights into the basic structure and essential issues of table placement methods in the context of big data processing infrastructures. In this paper, we present such a study. The basic structure of a data placement method consists of three core operations: row reordering, table partitioning, and data packing. All the existing placement methods are formed by these core operations with variations made by the three key factors: (1) the size of a horizontal logical subset of a table (or the size of a row group), (2) the function of mapping columns to column groups, and (3) the function of packing columns or column groups in a row group into physical blocks. We have designed and implemented a benchmarking tool to provide insights into how variations of each factor affect the I/O performance of reading data of a table stored by a table placement method. Based on our results, we give suggested actions to optimize table reading performance. Results from large-scale experiments have also confirmed that our findings are valid for production workloads. Finally, we present ORC File as a case study to show the effectiveness of our findings and suggested actions.
Yin Huai, Rubao Lee, Owen O'Malley, Xiaodong Zhang 0001
Proc. VLDB Endow.4