EDBT 2026 Demo / reviewers in the wild / expert
Wangda Zhang
dblp:124/6700
· DBLP profile ↗
8ranked-venue papers
6as first author
3since 2021 · last 2022
0000-0002-4965-8132ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 8 · 6 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Deploying a Steered Query Optimizer in Production at MicrosoftabstractModern analytical workloads are highly heterogeneous and massively complex, making generic out of the box query optimizers untenable for many customers and scenarios. As a result, it is important to specialize these optimizers to instances of the workloads. In this paper, we continue a recent line of work in steering a query optimizer towards better plans for a given workload, and make major strides in pushing previous research ideas to production deployment. Along the way we solve several operational challenges including, making steering actions more manageable, keeping the costs of steering within budget, and avoiding unexpected performance regressions in production. Our resulting system, QO-Advisor, essentially externalizes the query planner to a massive offline pipeline for better exploration and specialization. We discuss various aspects of our design and show detailed results over production SCOPE workloads at Microsoft, where the system is currently enabled by default. Wangda Zhang, Matteo Interlandi, Paul Mineiro, Shi Qiao 0001, Nasim Ghazanfari, Karlen Lie, Marc T. Friedman, Rafah Hosn, Hiren Patel, Alekh Jindal |
SIGMOD Conference | 1 |
| 2022 | Exploiting Data Skew for Improved Query PerformanceabstractAnalytic queries enable sophisticated large-scale data analysis within many commercial, scientific and medical domains today. Data skew is a ubiquitous feature of these real-world domains. In a retail database, some products are typically much more popular than others. In a text database, word frequencies follow a Zipf distribution with a small number of very common words, and a long tail of infrequent words. In a geographic database, some regions have much higher populations (and therefore data measurements) than others. Current systems do not make the most of caches for exploiting skew. In particular, a whole cache line may remain cache resident even though only a small part of the cache line corresponds to a popular data item. In this article, we propose a novel index structure for repositioning data items to concentrate popular items into the same cache lines. The net result is better spatial locality, and better utilization of limited cache resources. We develop a theoretical model for analyzing the cache utilization, and implement database operators that are efficient in the presence of skew. Our experimental evaluation on real and synthetic data shows that exploiting skew can significantly improve in-memory query performance. In some cases, our techniques can speed up queries by over an order of magnitude. Wangda Zhang, Kenneth A. Ross |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2021 | Adaptive Code Generation for Data-Intensive AnalyticsabstractModern database management systems employ sophisticated query optimization techniques that enable the generation of efficient plans for queries over very large data sets. A variety of other applications also process large data sets, but cannot leverage database-style query optimization for their code. We therefore identify an opportunity to enhance an open-source programming language compiler with database-style query optimization. Our system dynamically generates execution plans at query time, and runs those plans on chunks of data at a time. Based on feedback from earlier chunks, alternative plans might be used for later chunks. The compiler extension could be used for a variety of data-intensive applications, allowing all of them to benefit from this class of performance optimizations. Wangda Zhang, Junyoung Kim 0004, Kenneth A. Ross, Eric Sedlar, Lukas Stadler |
Proc. VLDB Endow. | 1 |
| 2020 | Permutation Index: Exploiting Data Skew for Improved Query PerformanceabstractAnalytic queries enable sophisticated large-scale data analysis within many commercial, scientific and medical domains today. Data skew is a ubiquitous feature of these real-world domains, but current systems do not make the most of caches for exploiting skew. In particular, a whole cache line may remain cache resident even though only a small part of the cache line corresponds to a popular data item. In this paper, we propose a novel index structure for repositioning data items to concentrate popular items into the same cache lines, resulting in better spatial locality, and better utilization of limited cache resources. We analyze cache behavior, and implement database operators that are efficient in the presence of skew. Experiments on real and synthetic data show that exploiting skew can significantly improve in-memory query performance. In some cases, our techniques can speed up queries by over an order of magnitude. Wangda Zhang, Kenneth A. Ross |
ICDE | 1 |
| 2020 | Efficient Search over Genomic Short Read DataabstractModern DNA sequencing technology produces large volumes of genome strings for various biological and medical applications. To mitigate the space overhead required for storing these genome data, previous research has developed compression schemes to save storage resources and to improve the locality for bioinformatics applications. These approaches, however, typically need significant additional processing to support efficient genome string search, an important step for downstream applications in a bioinformatics pipeline. In this paper, we propose to store raw DNA sequence data in a compressed but searchable format, enabling efficient string lookups using database indexes. To build this format, we partition genome strings by computing hash-based minimizers to group overlapping strings together. We carefully optimize our hash function to avoid large buckets caused by repetitive sequences. These buckets are then effectively compressed using local constructed references, and an index is built upon them to guide exact match lookups. Both data compression and string lookups exploit the enhanced locality as a result of the partitioning. We implement storage and search functions over the storage format as a multithreaded library, and perform extensive experiments on real genome sequencing data. The results show that our approach efficiently executes genome string lookups over highly compressed raw read data. Wangda Zhang, Mengdi Lin, Kenneth A. Ross |
SSDBM | 1 |
| 2018 | Distributed Joins and Data Placement for Minimal Network TrafficabstractNetwork communication is the slowest component of many operators in distributed parallel databases deployed for large-scale analytics. Whereas considerable work has focused on speeding up databases on modern hardware, communication reduction has received less attention. Existing parallel DBMSs rely on algorithms designed for disks with minor modifications for networks. A more complicated algorithm may burden the CPUs but could avoid redundant transfers of tuples across the network. We introduce track join, a new distributed join algorithm that minimizes network traffic by generating an optimal transfer schedule for each distinct join key. Track join extends the trade-off options between CPU and network. Track join explicitly detects and exploits locality, also allowing for advanced placement of tuples beyond hash partitioning on a single attribute. We propose a novel data placement algorithm based on track join that minimizes the total network cost of multiple joins across different dimensions in an analytical workload. Our evaluation shows that track join outperforms hash join on the most expensive queries of real workloads regarding both network traffic and execution time. Finally, we show that our data placement optimization approach is both robust and effective in minimizing the total network cost of joins in analytical workloads. Orestis Polychroniou, Wangda Zhang, Kenneth A. Ross |
ACM Trans. Database Syst. | 2 |
| 2015 | Discovering Meta-Paths in Large Heterogeneous Information NetworksabstractThe Heterogeneous Information Network (HIN) is a graph data model in which nodes and edges are annotated with class and relationship labels. Large and complex datasets, such as Yago or DBLP, can be modeled as HINs. Recent work has studied how to make use of these rich information sources. In particular, meta-paths, which represent sequences of node classes and edge types between two nodes in a HIN, have been proposed for such tasks as information retrieval, decision making, and product recommendation. Current methods assume meta-paths are found by domain experts. However, in a large and complex HIN, retrieving meta-paths manually can be tedious and difficult. We thus study how to discover meta-paths automatically. Specifically, users are asked to provide example pairs of nodes that exhibit high proximity. We then investigate how to generate meta-paths that can best explain the relationship between these node pairs. Since this problem is computationally intractable, we propose a greedy algorithm to select the most relevant meta-paths. We also present a data structure to enable efficient execution of this algorithm. We further incorporate hierarchical relationships among node classes in our solutions. Extensive experiments on real-world HIN show that our approach captures important meta-paths in an efficient and scalable manner. Changping Meng, Reynold Cheng, Silviu Maniu, Pierre Senellart, Wangda Zhang |
WWW | 5 |
| 2014 | Evaluating multi-way joins over discounted hitting timeabstractThe discounted hitting time (DHT), which is a random-walk similarity measure for graph node pairs, is useful in various applications, including link prediction, collaborative recommendation, and reputation ranking. We examine a novel query, called the multi-way join (or n-way join), on DHT scores. Given a graph and n sets of nodes, the n-way join retrieves a set of n-tuples with the k highest scores, according to some aggregation function of DHT values. This query enables analysis and prediction of complex relationship among n sets of nodes. Since an n-way join is expensive to compute, we develop the Partial Join algorithm (or PJ). This solution decomposes an n-way join into a number of top-m 2-way joins, and combines their results to construct the answer of the n-way join. Since PJ may necessitate the computation of top-(m+ 1) 2-way joins, we study an incremental solution, which allows the top-(m+ 1) 2-way join to be derived quickly from the top-m 2-way join results earlier computed. We further examine fast processing and pruning algorithms for 2-way joins. An extensive evaluation on three real datasets shows that PJ accurately evaluates n-way joins, and is four orders of magnitude faster than basic solutions. Wangda Zhang, Reynold Cheng, Ben Kao |
ICDE | 1 |