Wangchao Le

dblp:65/3302 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
1since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 8 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
9 papers
Query processing and optimization · 74% Data models and query languages · 8% Information retrieval · 6%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
Cloud and datacenter computing · 100%

Topics — the 23 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Query processing and optimization
query optimization
1.332022
Pipemizer: An Optimizer for Analytics Data Pipelines · Proc. VLDB Endow. 2022
Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Query processing and optimization
cost model
0.522020
Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020
Scalable Multi-query Optimization for SPARQL · ICDE 2012
Query processing and optimization › cost model
learned cost model
0.412020
Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020
Data models and query languages › semistructured data
RDF data
0.322014
Scalable Keyword Search on Large RDF Data · IEEE Trans. Knowl. Data Eng. 2014
Scalable Multi-query Optimization for SPARQL · ICDE 2012
Query processing and optimization
cardinality estimation
0.312018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Query processing and optimization › cardinality estimation
learned cardinality estimation
0.312018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Query processing and optimization › query optimization › learned query optimization
learned query optimizer
0.312018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Information retrieval
keyword search
0.212014
Scalable Keyword Search on Large RDF Data · IEEE Trans. Knowl. Data Eng. 2014
Information retrieval
text summarization
0.212014
Scalable Keyword Search on Large RDF Data · IEEE Trans. Knowl. Data Eng. 2014
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling
0.212022
Pipemizer: An Optimizer for Analytics Data Pipelines · Proc. VLDB Endow. 2022
Distributed and cloud data management › cloud database
database-as-a-service
0.112012
Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012
Query processing and optimization
multi-query optimization
0.112012
Scalable Multi-query Optimization for SPARQL · ICDE 2012
Query processing and optimization › secure query processing
outsourced database query authentication
0.112012
Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012
Query processing and optimization › query optimization › graph query optimization
SPARQL query optimization
0.112012
Scalable Multi-query Optimization for SPARQL · ICDE 2012
Graph data management › RDF data management
SPARQL query processing
0.112012
Scalable Multi-query Optimization for SPARQL · ICDE 2012
Cloud and datacenter computing
resource allocation
0.112020
Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020
Query processing and optimization
query rewriting
0.112011
Rewriting queries on SPARQL views · WWW 2011
Data models and query languages › RDF query language
SPARQL
0.112011
Rewriting queries on SPARQL views · WWW 2011
Query processing and optimization
top-k query processing
0.112010
Top-k queries on temporal data · VLDB J. 2010
Machine learning and data management
learned database components
0.112018
Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018
Cryptographic protocols and secure computation › verifiable computation
query result verification
0.012012
Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012
Cryptographic protocols and secure computation
verifiable computation
0.012012
Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012
Spatial and temporal data management
temporal databases
0.012010
Top-k queries on temporal data · VLDB J. 2010

Methods — techniques the papers use, named apart from their topics

machine learning · 1.5operator push-up · 1.1job split and merge · 1.1subgraph template learning · 0.7pruning · 0.2graph summarization · 0.2partitioning · 0.2load balancing · 0.2number theory · 0.1monte carlo verification · 0.1heuristic algorithm · 0.1common sub-structure discovery · 0.1
YearPublicationVenuePosition
2022 Pipemizer: An Optimizer for Analytics Data Pipelines
abstract
We demonstrate Pipemizer , an optimizer and recommender aimed at improving the performance of queries or jobs in pipelines. These job pipelines are ubiquitous in modern data analytics due to jobs reading output files written by other jobs. Given that more than 650k jobs run on Microsoft's SCOPE job service per day and about 70% have inter-job dependencies, identifying optimization opportunities across query jobs is of considerable interest to both cluster operators and users. Pipemizer addresses this need by providing recommendations to users, allowing users to understand their system, and facilitating automated application of recommendations. Pipemizer introduces novel optimizations that include holistic pipeline-aware statistics generation, inter-job operator push-up, and job split & merge. This demonstration showcases optimizations and recommendations generated by Pipemizer , enabling users to understand and optimize job pipelines.
Sunny Gakhar, Joyce Cahoon, Wangchao Le, Xiangnan Li, Kaushik Ravichandran 0002, Hiren Patel, Marc T. Friedman, Brandon Haynes, Shi Qiao 0001, Alekh Jindal, Jyoti Leeka
Proc. VLDB Endow.3
2020 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings
abstract
Query processing over big data is ubiquitous in modern clouds, where the system takes care of picking both the physical query execution plans and the resources needed to run those plans, using a cost-based query optimizer. A good cost model, therefore, is akin to better resource efficiency and lower operational costs. Unfortunately, the production workloads at Microsoft show that costs are very complex to model for big data systems. In this work, we investigate two key questions: (i) can we learn accurate cost models for big data systems, and (ii) can we integrate the learned models within the query optimizer. To answer these, we make three core contributions. First, we exploit workload patterns to learn a large number of individual cost models and combine them to achieve high accuracy and coverage over a long period. Second, we propose extensions to Cascades framework to pick optimal resources, i.e, number of containers, during query planning. And third, we integrate the learned cost models within the Cascade-style query optimizer of SCOPE at Microsoft. We evaluate the resulting system, Cleo, in a production environment using both production and TPC-H workloads. Our results show that the learned cost models are 2 to 3 orders of magnitude more accurate, and 20X more correlated with the actual runtimes, with a large majority (70%) of the plan changes leading to substantial improvements in latency as well as resource usage.
Tarique Siddiqui, Alekh Jindal, Shi Qiao 0001, Hiren Patel, Wangchao Le
SIGMOD Conference5
2018 Towards a Learning Optimizer for Shared Clouds
abstract
Query optimizers are notorious for inaccurate cost estimates, leading to poor performance. The root of the problem lies in inaccurate cardinality estimates, i.e., the size of intermediate (and final) results in a query plan. These estimates also determine the resources consumed in modern shared cloud infrastructures. In this paper, we present C ARD L EARNER , a machine learning based approach to learn cardinality models from previous job executions and use them to predict the cardinalities in future jobs. The key intuition in our approach is that shared cloud workloads are often recurring and overlapping in nature, and so we could learn cardinality models for overlapping subgraph templates. We discuss various learning approaches and show how learning a large number of smaller models results in high accuracy and explainability. We further present an exploration technique to avoid learning bias by considering alternate join orders and learning cardinality models over them. We describe the feedback loop to apply the learned models back to future job executions. Finally, we show a detailed evaluation of our models (up to 5 orders of magnitude less error), query plans (60% applicability), performance (up to 100% faster, 3x fewer resources), and exploration (optimal in few 10s of executions).
Chenggang Wu 0001, Alekh Jindal, Saeed Amizadeh, Hiren Patel, Wangchao Le, Shi Qiao 0001, Sriram Rao
Proc. VLDB Endow.5
2014 Scalable Keyword Search on Large RDF Data
abstract
Keyword search is a useful tool for exploring large RDF data sets. Existing techniques either rely on constructing a distance matrix for pruning the search space or building summaries from the RDF graphs for query processing. In this work, we show that existing techniques have serious limitations in dealing with realistic, large RDF data with tens of millions of triples. Furthermore, the existing summarization techniques may lead to incorrect/incomplete results. To address these issues, we propose an effective summarization algorithm to summarize the RDF data. Given a keyword query, the summaries lend significant pruning powers to exploratory keyword search and result in much better efficiency compared to previous works. Unlike existing techniques, our search algorithms always return correct results. Besides, the summaries we built can be updated incrementally and efficiently. Experiments on both benchmark and large real RDF data sets show that our techniques are scalable and efficient.
Wangchao Le, Feifei Li 0001, Anastasios Kementsietsidis, Songyun Duan
IEEE Trans. Knowl. Data Eng.1
2013 Optimal splitters for temporal and multi-version databases
abstract
Temporal and multi-version databases are ideal candidates for a distributed store, which offers large storage space, and parallel and distributed processing power from a cluster of (commodity) machines. A key challenge is to achieve a good load balancing algorithm for storage and processing of these data, which is done by partitioning the database. We introduce the concept of optimal splitters for temporal and multi-version databases, which induce a partition of the input data set, and guarantee that the size of the maximum bucket be minimized among all possible configurations, given a budget for the desired number of buckets. We design efficient methods for memory- and disk resident data respectively, and show that they significantly outperform competing baseline methods both theoretically and empirically on large real data sets.
Wangchao Le, Feifei Li 0001, Yufei Tao 0001, Robert Christensen
SIGMOD Conference1
2012 Scalable Multi-query Optimization for SPARQL
abstract
This paper revisits the classical problem of multi-query optimization in the context of RDF/SPARQL. We show that the techniques developed for relational and semi-structured data/query languages are hard, if not impossible, to be extended to account for RDF data model and graph query patterns expressed in SPARQL. In light of the NP-hardness of the multi-query optimization for SPARQL, we propose heuristic algorithms that partition the input batch of queries into groups such that each group of queries can be optimized together. An essential component of the optimization incorporates an efficient algorithm to discover the common sub-structures of multiple SPARQL queries and an effective cost model to compare candidate execution plans. Since our optimization techniques do not make any assumption about the underlying SPARQL query engine, they have the advantage of being portable across different RDF stores. The extensive experimental studies, performed on three popular RDF stores, show that the proposed techniques are effective, efficient and scalable.
Wangchao Le, Anastasios Kementsietsidis, Songyun Duan, Feifei Li 0001
ICDE1
2012 Query Access Assurance in Outsourced Databases
abstract
Query execution assurance is an important concept in defeating lazy servers in the database as a service model. We show that extending query execution assurance to outsourced databases with multiple data owners is highly inefficient. To cope with lazy servers in the distributed setting, we propose query access assurance (Qaa) that focuses on IO-bound queries. The goal in Qaa is to enable clients to verify that the server has honestly accessed all records that are necessary to compute the correct query answer, thus eliminating the incentives for the server to be lazy if the query cost is dominated by the IO cost in accessing these records. We formalize this concept for distributed databases, and present two efficient schemes that achieve Qaa with high success probabilities. The first scheme is simple to implement and deploy, but may incur excessive server to client communication cost and verification cost at the client side, when the query selectivity or the database size increases. The second scheme is more involved, but successfully addresses the limitation of the first scheme. Our design employs a few number theory techniques. Extensive experiments demonstrate the efficiency, effectiveness, and usefulness of our schemes.
Wangchao Le, Feifei Li 0001
IEEE Trans. Serv. Comput.1
2011 Rewriting queries on SPARQL views
abstract
The problem of answering SPARQL queries over virtual SPARQL views is commonly encountered in a number of settings, including while enforcing security policies to access RDF data, or when integrating RDF data from disparate sources. We approach this problem by rewriting SPARQL queries over the views to equivalent queries over the underlying RDF data, thus avoiding the costs entailed by view materialization and maintenance. We show that SPARQL query rewriting combines the most challenging aspects of rewriting for the relational and XML cases: like the relational case, SPARQL query rewriting requires synthesizing multiple views; like the XML case, the size of the rewritten query is exponential to the size of the query and the views. In this paper, we present the first native query rewriting algorithm for SPARQL. For an input SPARQL query over a set of virtual SPARQL views, the rewritten query resembles a union of conjunctive queries and can be of exponential size. We propose optimizations over the basic rewriting algorithm to (i) minimize each conjunctive query in the union; (ii) eliminate conjunctive queries with empty results from evaluation; and (iii) efficiently prune out big portions of the search space of empty rewritings. The experiments, performed on two RDF stores, show that our algorithms are scalable and independent of the underlying RDF stores. Furthermore, our optimizations have order of magnitude improvements over the basic rewriting algorithm in both the rewriting size and evaluation time.
Wangchao Le, Songyun Duan, Anastasios Kementsietsidis, Feifei Li 0001, Min Wang 0001
WWW1
2010 Top-k queries on temporal data
Feifei Li 0001, Ke Yi 0001, Wangchao Le
VLDB J.3