VLDB 2026 Research / reviewers in the wild / expert
Yin Huai
dblp:61/8080
· DBLP profile ↗
10ranked-venue papers
4as first author
0since 2021 · last 2016
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 6 · 2 first-authorSystems, architecture and hardware · 3 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 2Artificial intelligence and machine learning · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
5 papers |
Distributed and cloud data management · 30% Query processing and optimization · 20% Data models and query languages · 18% | |
| Computer architecture, parallel and distributed computing, and storage systems
4 papers |
Storage systems · 59% Performance modeling and evaluation · 29% GPUs and heterogeneous computing · 13% |
Topics — the 11 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Storage systems › data management › database storage
columnar storage |
0.3 | 2 | 2014 | Major technical advancements in apache hive · SIGMOD Conference 2014 RCFile: A fast and space-efficient data placement structure in MapReduce-based warehouse systems · ICDE 2011 |
Query processing and optimization › query optimization › query optimizer architecture
query optimizer extensibility |
0.2 | 1 | 2015 | Spark SQL: Relational Data Processing in Spark · SIGMOD Conference 2015 |
Data integration and cleaning
data warehouse |
0.2 | 1 | 2014 | Major technical advancements in apache hive · SIGMOD Conference 2014 |
Storage systems › data representation
file format |
0.2 | 1 | 2014 | Major technical advancements in apache hive · SIGMOD Conference 2014 |
Performance modeling and evaluation
benchmarking |
0.2 | 1 | 2013 | Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters · Proc. VLDB Endow. 2013 |
Storage systems
data placement |
0.2 | 1 | 2013 | Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters · Proc. VLDB Endow. 2013 |
Performance modeling and evaluation › benchmarking
i/o benchmarking |
0.2 | 1 | 2013 | Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in Clusters · Proc. VLDB Endow. 2013 |
GPUs and heterogeneous computing
CPU-GPU heterogeneous computing |
0.1 | 1 | 2012 | Accelerating Pathology Image Data Cross-Comparison on CPU-GPU Hybrid Systems · Proc. VLDB Endow. 2012 |
Distributed and cloud data management
data placement |
0.1 | 1 | 2011 | RCFile: A fast and space-efficient data placement structure in MapReduce-based warehouse systems · ICDE 2011 |
Distributed and cloud data management › mapreduce
mapreduce-based data warehouse |
0.1 | 1 | 2011 | RCFile: A fast and space-efficient data placement structure in MapReduce-based warehouse systems · ICDE 2011 |
Distributed and cloud data management › federated database
query federation |
0.1 | 1 | 2015 | Spark SQL: Relational Data Processing in Spark · SIGMOD Conference 2015 |
Methods — techniques the papers use, named apart from their topics
benchmarking · 0.4workload characterization · 0.3benchmarking tool · 0.3task migration · 0.3pipelining · 0.3row-store · 0.2hybrid-store · 0.2column-store · 0.2query optimization rules · 0.2code generation · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2016 | Spark-GPU: An accelerated in-memory data processing engine on clustersabstractApache Spark is an in-memory data processing system that supports both SQL queries and advanced analytics over large data sets. In this paper, we present our design and implementation of Spark-GPU that enables Spark to utilize GPU's massively parallel processing ability to achieve both high performance and high throughput. Spark-GPU transforms a general-purpose data processing system into a GPU-supported system by addressing several real-world technical challenges including minimizing internal and external data transfers, preparing a suitable data format and a batching mode for efficient GPU execution, and determining the suitability of workloads for GPU with a task scheduling capability between CPU and GPU. We have comprehensively evaluated Spark-GPU with a set of representative analytical workloads to show its effectiveness. Our results show that Spark-GPU improves the performance of machine learning workloads by up to 16.13x and the performance of SQL queries by up to 4.83x. Yuan Yuan 0014, Meisam Fathi Salmi, Yin Huai, Kaibo Wang, Rubao Lee, Xiaodong Zhang 0001 |
IEEE BigData | 3 |
| 2015 | SideWalk: A Facility of Lightweight Out-of-Band Communications for Augmenting Distributed Data Processing FlowsabstractThe foundation of a data processing engine running on a large cluster is its programming model that defines data processing operations and data movements. A special kind of communication activities that are not normally defined in the programming model but are often used in ad hoc ways in system development, is called out-of-band communications. The existing ad hoc solutions of out-of-band communications are often hard to reuse, error-prone, and not free from unwanted side effects. To address these issues, we have designed and implemented a standalone facility of out-of-band communications called SideWalk. With this facility, users can add out-of-band communication operations into their distributed data flows through a set of reusable APIs. These APIs have well defined semantics and thus, users' chances of writing error-prone programs with SideWalk are minimized. To prevent users from introducing unwanted side effects while using SideWalk, we prototype SideWalk to efficiently handle lightweight out-of-band communications and we restrict communication patterns that can be conducted through SideWalk without affecting the applicability of SideWalk on typical use cases. Our experimental results show that execution times of distributed data processing flows in a Hadoop environment with out-of-band communications implemented with SideWalk are reduced up to 1.53 times compared with that of distributed data processing flows with out-of-band communications implemented with a representative ad hoc solution. Yin Huai, Yuan Yuan 0014, Rubao Lee, Xiaodong Zhang 0001 |
CLUSTER | 1 |
| 2015 | Spark SQL: Relational Data Processing in SparkabstractSpark SQL is a new module in Apache Spark that integrates relational processing with Spark's functional programming API. Built on our experience with Shark, Spark SQL lets Spark programmers leverage the benefits of relational processing (e.g. declarative queries and optimized storage), and lets SQL users call complex analytics libraries in Spark (e.g. machine learning). Compared to previous systems, Spark SQL makes two main additions. First, it offers much tighter integration between relational and procedural processing, through a declarative DataFrame API that integrates with procedural Spark code. Second, it includes a highly extensible optimizer, Catalyst, built using features of the Scala programming language, that makes it easy to add composable rules, control code generation, and define extension points. Using Catalyst, we have built a variety of features (e.g. schema inference for JSON, machine learning types, and query federation to external databases) tailored for the complex needs of modern data analysis. We see Spark SQL as an evolution of both SQL-on-Spark and of Spark itself, offering richer APIs and optimizations while keeping the benefits of the Spark programming model. Michael Armbrust, Reynold Xin, Cheng Lian 0001, Yin Huai, Davies Liu, Joseph K. Bradley, Tomer Kaftan, Michael J. Franklin, Ali Ghodsi 0002, Matei Zaharia |
SIGMOD Conference | 4 |
| 2015 | Hetero-DB: Next Generation High-Performance Database Systems by Best Utilizing Heterogeneous Computing and Storage Resources
Kai Zhang 0006, Feng Chen 0005, Xiaoning Ding, Yin Huai, Rubao Lee, Kaibo Wang, Yuan Yuan 0014, Xiaodong Zhang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2014 | Major technical advancements in apache hiveabstractApache Hive is a widely used data warehouse system for Apache Hadoop, and has been adopted by many organizations for various big data analytics applications. Closely working with many users and organizations, we have identified several shortcomings of Hive in its file formats, query planning, and query execution, which are key factors determining the performance of Hive. In order to make Hive continuously satisfy the requests and requirements of processing increasingly high volumes data in a scalable and efficient way, we have set two goals related to storage and runtime performance in our efforts on advancing Hive. First, we aim to maximize the effective storage capacity and to accelerate data accesses to the data warehouse by updating the existing file formats. Second, we aim to significantly improve cluster resource utilization and runtime performance of Hive by developing a highly optimized query planner and a highly efficient query execution engine. In this paper, we present a community-based effort on technical advancements in Hive. Our performance evaluation shows that these advancements provide significant improvements on storage efficiency and query execution performance. This paper also shows how academic research lays a foundation for Hive to improve its daily operations. Yin Huai, Ashutosh Chauhan, Alan Gates, Günther Hagleitner, Eric N. Hanson, Owen O'Malley, Jitendra Pandey, Yuan Yuan 0014, Rubao Lee, Xiaodong Zhang 0001 |
SIGMOD Conference | 1 |
| 2013 | Understanding Insights into the Basic Structure and Essential Issues of Table Placement Methods in ClustersabstractA table placement method is a critical component in big data analytics on distributed systems. It determines the way how data values in a two-dimensional table are organized and stored in the underlying cluster. Based on Hadoop computing environments, several table placement methods have been proposed and implemented. However, a comprehensive and systematic study to understand, to compare, and to evaluate different table placement methods has not been done. Thus, it is highly desirable to gain important insights into the basic structure and essential issues of table placement methods in the context of big data processing infrastructures. In this paper, we present such a study. The basic structure of a data placement method consists of three core operations: row reordering, table partitioning, and data packing. All the existing placement methods are formed by these core operations with variations made by the three key factors: (1) the size of a horizontal logical subset of a table (or the size of a row group), (2) the function of mapping columns to column groups, and (3) the function of packing columns or column groups in a row group into physical blocks. We have designed and implemented a benchmarking tool to provide insights into how variations of each factor affect the I/O performance of reading data of a table stored by a table placement method. Based on our results, we give suggested actions to optimize table reading performance. Results from large-scale experiments have also confirmed that our findings are valid for production workloads. Finally, we present ORC File as a case study to show the effectiveness of our findings and suggested actions. Yin Huai, Rubao Lee, Owen O'Malley, Xiaodong Zhang 0001 |
Proc. VLDB Endow. | 1 |
| 2012 | Accelerating Pathology Image Data Cross-Comparison on CPU-GPU Hybrid SystemsabstractAs an important application of spatial databases in pathology imaging analysis, cross-comparing the spatial boundaries of a huge amount of segmented micro-anatomic objects demands extremely data- and compute-intensive operations, requiring high throughput at an affordable cost. However, the performance of spatial database systems has not been satisfactory since their implementations of spatial operations cannot fully utilize the power of modern parallel hardware. In this paper, we provide a customized software solution that exploits GPUs and multi-core CPUs to accelerate spatial cross-comparison in a cost-effective way. Our solution consists of an efficient GPU algorithm and a pipelined system framework with task migration support. Extensive experiments with real-world data sets demonstrate the effectiveness of our solution, which improves the performance of spatial cross-comparison by over 18 times compared with a parallelized spatial database approach. Kaibo Wang, Yin Huai, Rubao Lee, Fusheng Wang 0001, Xiaodong Zhang 0001, Joel H. Saltz |
Proc. VLDB Endow. | 2 |
| 2011 | DOT: a matrix model for analyzing, optimizing and deploying software for big data analytics in distributed systemsabstractTraditional parallel processing models, such as BSP, are "scale up" based, aiming to achieve high performance by increasing computing power, interconnection network bandwidth, and memory/storage capacity within dedicated systems, while big data analytics tasks aiming for high throughput demand that large distributed systems "scale out" by continuously adding computing and storage resources through networks. Each one of the "scale up" model and "scale out" model has a different set of performance requirements and system bottlenecks. In this paper, we develop a general model that abstracts critical computation and communication behavior and computation-communication interactions for big data analytics in a scalable and fault-tolerant manner. Our model is called DOT, represented by three matrices for data sets (D), concurrent data processing operations (O), and data transformations (T), respectively. With the DOT model, any big data analytics job execution in various software frameworks can be represented by a specific or non-specific number of elementary/composite DOT blocks, each of which performs operations on the data sets, stores intermediate results, makes necessary data transfers, and performs data transformations in the end. The DOT model achieves the goals of scalability and fault-tolerance by enforcing a data-dependency-free relationship among concurrent tasks. Under the DOT model, we provide a set of optimization guidelines, which are framework and implementation independent, and applicable to a wide variety of big data analytics jobs. Finally, we demonstrate the effectiveness of the DOT model through several case studies. Yin Huai, Rubao Lee, Simon Zhang, Cathy H. Xia, Xiaodong Zhang 0001 |
SoCC | 1 |
| 2011 | YSmart: Yet Another SQL-to-MapReduce TranslatorabstractMapReduce has become an effective approach to big data analytics in large cluster systems, where SQL-like queries play important roles to interface between users and systems. However, based on our Facebook daily operation results, certain types of queries are executed at an unacceptable low speed by Hive (a production SQL-to-MapReduce translator). In this paper, we demonstrate that existing SQL-to-MapReduce translators that operate in a one-operation-to-one-job mode and do not consider query correlations cannot generate high-performance MapReduce programs for certain queries, due to the mismatch between complex SQL structures and simple MapReduce framework. We propose and develop a system called Y Smart, a correlation aware SQL-to-MapReduce translator. Y Smart applies a set of rules to use the minimal number of MapReduce jobs to execute multiple correlated operations in a complex query. Y Smart can significantly reduce redundant computations, I/O operations and network transfers compared to existing translators. We have implemented Y Smart with intensive evaluation for complex queries on two Amazon EC2 clusters and one Facebook production cluster. The results show that Y Smart can outperform Hive and Pig, two widely used SQL-to-MapReduce translators, by more than four times for query execution. Rubao Lee, Yin Huai, Fusheng Wang 0001, Yongqiang He, Xiaodong Zhang 0001 |
ICDCS | 3 |
| 2011 | RCFile: A fast and space-efficient data placement structure in MapReduce-based warehouse systemsabstractMapReduce-based data warehouse systems are playing important roles of supporting big data analytics to understand quickly the dynamics of user behavior trends and their needs in typical Web service providers and social network sites (e.g., Facebook). In such a system, the data placement structure is a critical factor that can affect the warehouse performance in a fundamental way. Based on our observations and analysis of Facebook production systems, we have characterized four requirements for the data placement structure: (1) fast data loading, (2) fast query processing, (3) highly efficient storage space utilization, and (4) strong adaptivity to highly dynamic workload patterns. We have examined three commonly accepted data placement structures in conventional databases, namely row-stores, column-stores, and hybrid-stores in the context of large data analysis using MapReduce. We show that they are not very suitable for big data processing in distributed systems. In this paper, we present a big data placement structure called RCFile (Record Columnar File) and its implementation in the Hadoop system. With intensive experiments, we show the effectiveness of RCFile in satisfying the four requirements. RCFile has been chosen in Facebook data warehouse system as the default option. It has also been adopted by Hive and Pig, the two most widely used data analysis systems developed in Facebook and Yahoo! Yongqiang He, Rubao Lee, Yin Huai, Zheng Shao, Namit Jain, Xiaodong Zhang 0001 |
ICDE | 3 |