EDBT 2026 Demo / reviewers in the wild / expert
Wangchao Le
dblp:65/3302
· DBLP profile ↗
9ranked-venue papers
5as first author
1since 2021 · last 2022
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 4 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
9 papers |
Query processing and optimization · 74% Data models and query languages · 8% Information retrieval · 6% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
Cloud and datacenter computing · 100% |
Topics — the 23 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
query optimization |
1.3 | 3 | 2022 | Pipemizer: An Optimizer for Analytics Data Pipelines · Proc. VLDB Endow. 2022 Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020 Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018 |
Query processing and optimization
cost model |
0.5 | 2 | 2020 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020 Scalable Multi-query Optimization for SPARQL · ICDE 2012 |
Query processing and optimization › cost model
learned cost model |
0.4 | 1 | 2020 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020 |
Data models and query languages › semistructured data
RDF data |
0.3 | 2 | 2014 | Scalable Keyword Search on Large RDF Data · IEEE Trans. Knowl. Data Eng. 2014 Scalable Multi-query Optimization for SPARQL · ICDE 2012 |
Query processing and optimization
cardinality estimation |
0.3 | 1 | 2018 | Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018 |
Query processing and optimization › cardinality estimation
learned cardinality estimation |
0.3 | 1 | 2018 | Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018 |
Query processing and optimization › query optimization › learned query optimization
learned query optimizer |
0.3 | 1 | 2018 | Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018 |
Information retrieval
keyword search |
0.2 | 1 | 2014 | Scalable Keyword Search on Large RDF Data · IEEE Trans. Knowl. Data Eng. 2014 |
Information retrieval
text summarization |
0.2 | 1 | 2014 | Scalable Keyword Search on Large RDF Data · IEEE Trans. Knowl. Data Eng. 2014 |
Cloud and datacenter computing › cluster resource management and scheduling
cluster scheduling |
0.2 | 1 | 2022 | Pipemizer: An Optimizer for Analytics Data Pipelines · Proc. VLDB Endow. 2022 |
Distributed and cloud data management › cloud database
database-as-a-service |
0.1 | 1 | 2012 | Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012 |
Query processing and optimization
multi-query optimization |
0.1 | 1 | 2012 | Scalable Multi-query Optimization for SPARQL · ICDE 2012 |
Query processing and optimization › secure query processing
outsourced database query authentication |
0.1 | 1 | 2012 | Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012 |
Query processing and optimization › query optimization › graph query optimization
SPARQL query optimization |
0.1 | 1 | 2012 | Scalable Multi-query Optimization for SPARQL · ICDE 2012 |
Graph data management › RDF data management
SPARQL query processing |
0.1 | 1 | 2012 | Scalable Multi-query Optimization for SPARQL · ICDE 2012 |
Cloud and datacenter computing
resource allocation |
0.1 | 1 | 2020 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our Findings · SIGMOD Conference 2020 |
Query processing and optimization
query rewriting |
0.1 | 1 | 2011 | Rewriting queries on SPARQL views · WWW 2011 |
Data models and query languages › RDF query language
SPARQL |
0.1 | 1 | 2011 | Rewriting queries on SPARQL views · WWW 2011 |
Query processing and optimization
top-k query processing |
0.1 | 1 | 2010 | Top-k queries on temporal data · VLDB J. 2010 |
Machine learning and data management
learned database components |
0.1 | 1 | 2018 | Towards a Learning Optimizer for Shared Clouds · Proc. VLDB Endow. 2018 |
Cryptographic protocols and secure computation › verifiable computation
query result verification |
0.0 | 1 | 2012 | Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012 |
Cryptographic protocols and secure computation
verifiable computation |
0.0 | 1 | 2012 | Query Access Assurance in Outsourced Databases · IEEE Trans. Serv. Comput. 2012 |
Spatial and temporal data management
temporal databases |
0.0 | 1 | 2010 | Top-k queries on temporal data · VLDB J. 2010 |
Methods — techniques the papers use, named apart from their topics
machine learning · 1.5operator push-up · 1.1job split and merge · 1.1subgraph template learning · 0.7pruning · 0.2graph summarization · 0.2partitioning · 0.2load balancing · 0.2number theory · 0.1monte carlo verification · 0.1heuristic algorithm · 0.1common sub-structure discovery · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Pipemizer: An Optimizer for Analytics Data PipelinesabstractWe demonstrate Pipemizer , an optimizer and recommender aimed at improving the performance of queries or jobs in pipelines. These job pipelines are ubiquitous in modern data analytics due to jobs reading output files written by other jobs. Given that more than 650k jobs run on Microsoft's SCOPE job service per day and about 70% have inter-job dependencies, identifying optimization opportunities across query jobs is of considerable interest to both cluster operators and users. Pipemizer addresses this need by providing recommendations to users, allowing users to understand their system, and facilitating automated application of recommendations. Pipemizer introduces novel optimizations that include holistic pipeline-aware statistics generation, inter-job operator push-up, and job split & merge. This demonstration showcases optimizations and recommendations generated by Pipemizer , enabling users to understand and optimize job pipelines. Sunny Gakhar, Joyce Cahoon, Wangchao Le, Xiangnan Li, Kaushik Ravichandran 0002, Hiren Patel, Marc T. Friedman, Brandon Haynes, Shi Qiao 0001, Alekh Jindal, Jyoti Leeka |
Proc. VLDB Endow. | 3 |
| 2020 | Cost Models for Big Data Query Processing: Learning, Retrofitting, and Our FindingsabstractQuery processing over big data is ubiquitous in modern clouds, where the system takes care of picking both the physical query execution plans and the resources needed to run those plans, using a cost-based query optimizer. A good cost model, therefore, is akin to better resource efficiency and lower operational costs. Unfortunately, the production workloads at Microsoft show that costs are very complex to model for big data systems. In this work, we investigate two key questions: (i) can we learn accurate cost models for big data systems, and (ii) can we integrate the learned models within the query optimizer. To answer these, we make three core contributions. First, we exploit workload patterns to learn a large number of individual cost models and combine them to achieve high accuracy and coverage over a long period. Second, we propose extensions to Cascades framework to pick optimal resources, i.e, number of containers, during query planning. And third, we integrate the learned cost models within the Cascade-style query optimizer of SCOPE at Microsoft. We evaluate the resulting system, Cleo, in a production environment using both production and TPC-H workloads. Our results show that the learned cost models are 2 to 3 orders of magnitude more accurate, and 20X more correlated with the actual runtimes, with a large majority (70%) of the plan changes leading to substantial improvements in latency as well as resource usage. Tarique Siddiqui, Alekh Jindal, Shi Qiao 0001, Hiren Patel, Wangchao Le |
SIGMOD Conference | 5 |
| 2018 | Towards a Learning Optimizer for Shared CloudsabstractQuery optimizers are notorious for inaccurate cost estimates, leading to poor performance. The root of the problem lies in inaccurate cardinality estimates, i.e., the size of intermediate (and final) results in a query plan. These estimates also determine the resources consumed in modern shared cloud infrastructures. In this paper, we present C ARD L EARNER , a machine learning based approach to learn cardinality models from previous job executions and use them to predict the cardinalities in future jobs. The key intuition in our approach is that shared cloud workloads are often recurring and overlapping in nature, and so we could learn cardinality models for overlapping subgraph templates. We discuss various learning approaches and show how learning a large number of smaller models results in high accuracy and explainability. We further present an exploration technique to avoid learning bias by considering alternate join orders and learning cardinality models over them. We describe the feedback loop to apply the learned models back to future job executions. Finally, we show a detailed evaluation of our models (up to 5 orders of magnitude less error), query plans (60% applicability), performance (up to 100% faster, 3x fewer resources), and exploration (optimal in few 10s of executions). Chenggang Wu 0001, Alekh Jindal, Saeed Amizadeh, Hiren Patel, Wangchao Le, Shi Qiao 0001, Sriram Rao |
Proc. VLDB Endow. | 5 |
| 2014 | Scalable Keyword Search on Large RDF DataabstractKeyword search is a useful tool for exploring large RDF data sets. Existing techniques either rely on constructing a distance matrix for pruning the search space or building summaries from the RDF graphs for query processing. In this work, we show that existing techniques have serious limitations in dealing with realistic, large RDF data with tens of millions of triples. Furthermore, the existing summarization techniques may lead to incorrect/incomplete results. To address these issues, we propose an effective summarization algorithm to summarize the RDF data. Given a keyword query, the summaries lend significant pruning powers to exploratory keyword search and result in much better efficiency compared to previous works. Unlike existing techniques, our search algorithms always return correct results. Besides, the summaries we built can be updated incrementally and efficiently. Experiments on both benchmark and large real RDF data sets show that our techniques are scalable and efficient. Wangchao Le, Feifei Li 0001, Anastasios Kementsietsidis, Songyun Duan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2013 | Optimal splitters for temporal and multi-version databasesabstractTemporal and multi-version databases are ideal candidates for a distributed store, which offers large storage space, and parallel and distributed processing power from a cluster of (commodity) machines. A key challenge is to achieve a good load balancing algorithm for storage and processing of these data, which is done by partitioning the database. We introduce the concept of optimal splitters for temporal and multi-version databases, which induce a partition of the input data set, and guarantee that the size of the maximum bucket be minimized among all possible configurations, given a budget for the desired number of buckets. We design efficient methods for memory- and disk resident data respectively, and show that they significantly outperform competing baseline methods both theoretically and empirically on large real data sets. Wangchao Le, Feifei Li 0001, Yufei Tao 0001, Robert Christensen |
SIGMOD Conference | 1 |
| 2012 | Scalable Multi-query Optimization for SPARQLabstractThis paper revisits the classical problem of multi-query optimization in the context of RDF/SPARQL. We show that the techniques developed for relational and semi-structured data/query languages are hard, if not impossible, to be extended to account for RDF data model and graph query patterns expressed in SPARQL. In light of the NP-hardness of the multi-query optimization for SPARQL, we propose heuristic algorithms that partition the input batch of queries into groups such that each group of queries can be optimized together. An essential component of the optimization incorporates an efficient algorithm to discover the common sub-structures of multiple SPARQL queries and an effective cost model to compare candidate execution plans. Since our optimization techniques do not make any assumption about the underlying SPARQL query engine, they have the advantage of being portable across different RDF stores. The extensive experimental studies, performed on three popular RDF stores, show that the proposed techniques are effective, efficient and scalable. Wangchao Le, Anastasios Kementsietsidis, Songyun Duan, Feifei Li 0001 |
ICDE | 1 |
| 2012 | Query Access Assurance in Outsourced DatabasesabstractQuery execution assurance is an important concept in defeating lazy servers in the database as a service model. We show that extending query execution assurance to outsourced databases with multiple data owners is highly inefficient. To cope with lazy servers in the distributed setting, we propose query access assurance (Qaa) that focuses on IO-bound queries. The goal in Qaa is to enable clients to verify that the server has honestly accessed all records that are necessary to compute the correct query answer, thus eliminating the incentives for the server to be lazy if the query cost is dominated by the IO cost in accessing these records. We formalize this concept for distributed databases, and present two efficient schemes that achieve Qaa with high success probabilities. The first scheme is simple to implement and deploy, but may incur excessive server to client communication cost and verification cost at the client side, when the query selectivity or the database size increases. The second scheme is more involved, but successfully addresses the limitation of the first scheme. Our design employs a few number theory techniques. Extensive experiments demonstrate the efficiency, effectiveness, and usefulness of our schemes. Wangchao Le, Feifei Li 0001 |
IEEE Trans. Serv. Comput. | 1 |
| 2011 | Rewriting queries on SPARQL viewsabstractThe problem of answering SPARQL queries over virtual SPARQL views is commonly encountered in a number of settings, including while enforcing security policies to access RDF data, or when integrating RDF data from disparate sources. We approach this problem by rewriting SPARQL queries over the views to equivalent queries over the underlying RDF data, thus avoiding the costs entailed by view materialization and maintenance. We show that SPARQL query rewriting combines the most challenging aspects of rewriting for the relational and XML cases: like the relational case, SPARQL query rewriting requires synthesizing multiple views; like the XML case, the size of the rewritten query is exponential to the size of the query and the views. In this paper, we present the first native query rewriting algorithm for SPARQL. For an input SPARQL query over a set of virtual SPARQL views, the rewritten query resembles a union of conjunctive queries and can be of exponential size. We propose optimizations over the basic rewriting algorithm to (i) minimize each conjunctive query in the union; (ii) eliminate conjunctive queries with empty results from evaluation; and (iii) efficiently prune out big portions of the search space of empty rewritings. The experiments, performed on two RDF stores, show that our algorithms are scalable and independent of the underlying RDF stores. Furthermore, our optimizations have order of magnitude improvements over the basic rewriting algorithm in both the rewriting size and evaluation time. Wangchao Le, Songyun Duan, Anastasios Kementsietsidis, Feifei Li 0001, Min Wang 0001 |
WWW | 1 |
| 2010 | Top-k queries on temporal data
Feifei Li 0001, Ke Yi 0001, Wangchao Le |
VLDB J. | 3 |