EDBT 2026 Demo / reviewers in the wild / expert
Zeyuan Shang
dblp:189/2405
· DBLP profile ↗
10ranked-venue papers
6as first author
2since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Databases, data management, data science and information retrieval · 9 · 6 first-author · 1 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Databases, data mining, and information retrieval
10 papers |
Query processing and optimization · 39% Information retrieval · 15% Distributed and cloud data management · 14% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Query processing and optimization
similarity join |
1.2 | 4 | 2019 | Balance-Aware Distributed String Similarity-Based Query Processing System · Proc. VLDB Endow. 2019 Dima: A Distributed In-Memory Similarity-Based Query Processing System · Proc. VLDB Endow. 2017 K-Join: Knowledge-Aware Similarity Join · ICDE 2017 |
Information retrieval
retrieval models |
0.9 | 1 | 2025 | Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models · ICLR 2025 |
Machine learning and data management › model evaluation
retrieval model evaluation |
0.9 | 1 | 2025 | Beyond Content Relevance: Evaluating Instruction Following in Retrieval Models · ICLR 2025 |
Distributed and cloud data management
distributed query processing |
0.7 | 2 | 2019 | Balance-Aware Distributed String Similarity-Based Query Processing System · Proc. VLDB Endow. 2019 Dima: A Distributed In-Memory Similarity-Based Query Processing System · Proc. VLDB Endow. 2017 |
Query processing and optimization
similarity query processing |
0.7 | 2 | 2019 | Balance-Aware Distributed String Similarity-Based Query Processing System · Proc. VLDB Endow. 2019 Dima: A Distributed In-Memory Similarity-Based Query Processing System · Proc. VLDB Endow. 2017 |
Distributed and cloud data management › distributed analytics
distributed in-memory analytics |
0.7 | 2 | 2018 | DITA: A Distributed In-Memory Trajectory Analytics System · SIGMOD Conference 2018 DITA: Distributed In-Memory Trajectory Analytics · SIGMOD Conference 2018 |
Spatial and temporal data management
trajectory data management |
0.7 | 2 | 2018 | DITA: A Distributed In-Memory Trajectory Analytics System · SIGMOD Conference 2018 DITA: Distributed In-Memory Trajectory Analytics · SIGMOD Conference 2018 |
Query processing and optimization
approximate query processing |
0.5 | 1 | 2021 | Davos: A System for Interactive Data-Driven Decision Making · Proc. VLDB Endow. 2021 |
Query processing and optimization › approximate query processing
sampling-based approximate query processing |
0.5 | 1 | 2021 | Davos: A System for Interactive Data-Driven Decision Making · Proc. VLDB Endow. 2021 |
Query processing and optimization
cardinality estimation |
0.4 | 1 | 2019 | How I Learned to Stop Worrying and Love Re-optimization · ICDE 2019 |
Data integration and cleaning
data preprocessing |
0.4 | 1 | 2019 | Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019 |
Machine learning and data management
machine learning pipeline |
0.4 | 1 | 2019 | Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019 |
Query processing and optimization › adaptive query processing
query re-optimization |
0.4 | 1 | 2019 | How I Learned to Stop Worrying and Love Re-optimization · ICDE 2019 |
Information retrieval
similarity measure |
0.3 | 1 | 2017 | K-Join: Knowledge-Aware Similarity Join · ICDE 2017 |
Information retrieval
similarity search |
0.3 | 1 | 2017 | Dima: A Distributed In-Memory Similarity-Based Query Processing System · Proc. VLDB Endow. 2017 |
Data integration and cleaning
entity matching |
0.2 | 1 | 2016 | K-Join: Knowledge-Aware Similarity Join · IEEE Trans. Knowl. Data Eng. 2016 |
Spatial and temporal data management › spatial query processing
filtering-and-verification |
0.2 | 2 | 2018 | DITA: A Distributed In-Memory Trajectory Analytics System · SIGMOD Conference 2018 DITA: Distributed In-Memory Trajectory Analytics · SIGMOD Conference 2018 |
Human-AI interaction
interactive machine learning |
0.1 | 1 | 2019 | Democratizing Data Science through Interactive Curation of ML Pipelines · SIGMOD Conference 2019 |
Methods — techniques the papers use, named apart from their topics
re-ranking · 0.9LLM-based dense retrieval · 0.9signature-based indexing · 0.7query optimization · 0.7partitioning · 0.7local index · 0.7global index · 0.7cost-based load balancing · 0.7signature-based filtering · 0.5filter-and-verification · 0.5model selection · 0.4hyperparameter tuning · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Beyond Content Relevance: Evaluating Instruction Following in Retrieval ModelsabstractInstruction-following capabilities in large language models (LLMs) have progressed significantly, enabling more complex user interactions through detailed prompts. However, retrieval systems have not matched these advances, most of them still relies on traditional lexical and semantic matching techniques that fail to fully capture user intent. Recent efforts have introduced instruction-aware retrieval models, but these primarily focus on intrinsic content relevance, which neglects the importance of customized preferences for broader document-level attributes. This study evaluates the instruction-following capabilities of various retrieval models beyond content relevance, including LLM-based dense retrieval and reranking models. We develop InfoSearch, a novel retrieval evaluation benchmark spanning six document-level attributes: Audience, Keyword, Format, Language, Length, and Source, and introduce novel metrics -- Strict Instruction Compliance Ratio (SICR) and Weighted Instruction Sensitivity Evaluation (WISE) to accurately assess the models' responsiveness to instructions. Our findings indicate that although fine-tuning models on instruction-aware retrieval datasets and increasing model size enhance performance, most models still fall short of instruction compliance. We release our dataset and code on https://github.com/EIT-NLP/InfoSearch. Jianqun Zhou, Yuanlei Zheng, Zeyuan Shang, Wei Zhang 0185, Xiaoyu Shen 0001 |
ICLR | 5 |
| 2021 | Davos: A System for Interactive Data-Driven Decision MakingabstractRecently, a new horizon in data analytics, prescriptive analytics, is becoming more and more important to make data-driven decisions. As opposed to the progress of democratizing data acquisition and access, making data-driven decisions remains a significant challenge for people without technical expertise. In this regard, existing tools for data analytics which were designed decades ago still present a high bar for domain experts, and removing this bar requires a fundamental rethinking of both interface and backend. At Einblick, an MIT/Brown spin-off based on the Northstar project, we have been building the next generation analytics tool in the last few years. To overcome the shortcomings of existing processing engines, we propose Davos , Einblick's novel backend. Davos combines aspects of progressive computation, approximate query processing and sampling, with a specific focus on supporting user-defined operations. Moreover, Davos optimizes multi-tenant scenarios to promote collaboration. Both empirical evaluation and user study verify that Davos can greatly empower data analytics for new needs. Zeyuan Shang, Emanuel Zgraggen, Benedetto Buratti, Philipp Eichmann, Navid Karimeddiny, Charlie Meyer, Wesley Runnels, Tim Kraska |
Proc. VLDB Endow. | 1 |
| 2019 | How I Learned to Stop Worrying and Love Re-optimizationabstractCost-based query optimizers remain one of the most important components of database management systems for analytic workloads. Though modern optimizers select plans close to optimal performance in the common case, a small number of queries are an order of magnitude slower than they could be. In this paper we investigate why this is still the case, despite decades of improvements to cost models, plan enumeration, and cardinality estimation. We demonstrate why we believe that a re-optimization mechanism is likely the most cost-effective way to improve end-to-end query performance. We find that even a simple re-optimization scheme can improve the latency of many poorly performing queries. We demonstrate that re-optimization improves the end-to-end latency of the top 20 longest running queries in the Join Order Benchmark by 27%, realizing most of the benefit of perfect cardinality estimation. Matthew Perron, Zeyuan Shang, Tim Kraska, Michael Stonebraker |
ICDE | 2 |
| 2019 | Democratizing Data Science through Interactive Curation of ML PipelinesabstractStatistical knowledge and domain expertise are key to extract actionable insights out of data, yet such skills rarely coexist together. In Machine Learning, high-quality results are only attainable via mindful data preprocessing, hyperparameter tuning and model selection. Domain experts are often overwhelmed by such complexity, de-facto inhibiting a wider adoption of ML techniques in other fields. Existing libraries that claim to solve this problem, still require well-trained practitioners. Those frameworks involve heavy data preparation steps and are often too slow for interactive feedback from the user, severely limiting the scope of such systems. Zeyuan Shang, Emanuel Zgraggen, Benedetto Buratti, Ferdinand Kossmann, Philipp Eichmann, Yeounoh Chung, Carsten Binnig, Eli Upfal, Tim Kraska |
SIGMOD Conference | 1 |
| 2019 | Balance-Aware Distributed String Similarity-Based Query Processing SystemabstractData analysts spend more than 80% of time on data cleaning and integration in the whole process of data analytics due to data errors and inconsistencies. Similarity-based query processing is an important way to tolerate the errors and inconsistencies. However, similarity-based query processing is rather costly and traditional database cannot afford such expensive requirement. In this paper, we develop a distributed in-memory similarity-based query processing system called Dima. Dima supports four core similarity operations, i.e., similarity selection, similarity join, top- k selection and top- k join. Dima extends SQL for users to easily invoke these similarity-based operations in their data analysis tasks. To avoid expensive data transmission in a distributed environment, we propose balance-aware signatures where two records are similar if they share common signatures, and we can adaptively select the signatures to balance the workload. Dima builds signature-based global indexes and local indexes to support similarity operations. Since Spark is one of the widely adopted distributed in-memory computing systems, we have seamlessly integrated Dima into Spark and developed effective query optimization techniques in Spark. To the best of our knowledge, this is the first full-fledged distributed in-memory system that can support complex similarity-based query processing on large-scale datasets. We have conducted extensive experiments on four real-world datasets. Experimental results show that Dima outperforms state-of-the-art studies by 1--3 orders of magnitude and has good scalability. Ji Sun 0001, Zeyuan Shang, Guoliang Li 0001, Zhifeng Bao, Dong Deng 0001 |
Proc. VLDB Endow. | 2 |
| 2018 | DITA: Distributed In-Memory Trajectory AnalyticsabstractTrajectory analytics can benefit many real-world applications, e.g., frequent trajectory based navigation systems, road planning, car pooling, and transportation optimizations. Existing algorithms focus on optimizing this problem in a single machine. However, the amount of trajectories exceeds the storage and processing capability of a single machine, and it calls for large-scale trajectory analytics in distributed environments. The distributed trajectory analytics faces challenges of data locality aware partitioning, load balance, easy-to-use interface, and versatility to support various trajectory similarity functions. To address these challenges, we propose a distributed in-memory trajectory analytics system DITA. We propose an effective partitioning method, global index and local index, to address the data locality problem. We devise cost-based techniques to balance the workload. We develop a filter-verification framework to improve the performance. Moreover, DITA can support most of existing similarity functions to quantify the similarity between trajectories. We integrate our framework seamlessly into Spark SQL, and make it support SQL and DataFrame API interfaces. We have conducted extensive experiments on real world datasets, and experimental results show that DITA outperforms existing distributed trajectory similarity search and join approaches significantly. Zeyuan Shang, Guoliang Li 0001, Zhifeng Bao |
SIGMOD Conference | 1 |
| 2018 | DITA: A Distributed In-Memory Trajectory Analytics SystemabstractTrajectory analytics can benefit many real-world applications, e.g., frequent trajectory based navigation systems, road planning, car pooling, and transportation optimizations. In this paper, we demonstrate a distributed in-memory trajectory analytics system DITA to support large-scale trajectory data analytics. DITA exhibit three unique features. First, DITA supports threshold-based and KNN-based trajectory similarity search and join operations, as well as range queries (i.e., space and time). Second, DITA is versatile to support most existing similarity functions to cater for different analytic purposes and scenarios. Last, DITA is seamlessly integrated into Spark SQL to support easy-to-use SQL and DataFrame API interfaces. Technically, DITA proposes an effective partitioning method, global index and local index, to address the data locality problem. It also devises cost-based techniques to balance the workload, and develops a filter-verification framework for efficient and scalable search and join. Zeyuan Shang, Guoliang Li 0001, Zhifeng Bao |
SIGMOD Conference | 1 |
| 2017 | K-Join: Knowledge-Aware Similarity JoinabstractSimilarity join is a fundamental operation in data cleaning and integration. Existing similarity-join methods utilize the string similarity to quantify the relevance but neglect the knowledge behind the data, which plays an important role in understanding the data. Thanks to public knowledge bases, e.g., Freebase and Yago, we have an opportunity to use the knowledge to improve similarity join. To address this problem, we study knowledge-aware similarity join, which, given a knowledge hierarchy and two collections of objects (e.g., documents), finds all knowledge-aware similar object pairs. To the best of our knowledge, this is the first study on knowledge-aware similarity join. There are two main challenges. The first is how to quantify the knowledge-aware similarity. The second is how to efficiently identify the similar pairs. To address these challenges, we first propose a new similarity metric to quantify the knowledgeaware similarity using the knowledge hierarchy. We then devise a filter-and-verification framework to efficiently identify the similar pairs. We propose effective signature-based filtering techniques to prune large numbers of dissimilar pairs and develop efficient verification algorithms to verify the candidates that are not pruned in the filter step. Experimental results on real-world datasets show that our method significantly outperforms baseline algorithms in terms of both efficiency and effectiveness. Zeyuan Shang, Yaxiao Liu, Guoliang Li 0001, Jianhua Feng |
ICDE | 1 |
| 2017 | Dima: A Distributed In-Memory Similarity-Based Query Processing SystemabstractData analysts in industries spend more than 80% of time on data cleaning and integration in the whole process of data analytics due to data errors and inconsistencies. It calls for effective query processing techniques to tolerate the errors and inconsistencies. In this paper, we develop a distributed in-memory similarity-based query processing system called Dima. Dima supports two core similarity-based query operations, i.e., similarity search and similarity join. Dima extends the SQL programming interface for users to easily invoke these two operations in their data analysis jobs. To avoid expensive data transformation in a distributed environment, we design selectable signatures where two records approximately match if they share common signatures. More importantly, we can adaptively select the signatures to balance the workload. Dima builds signature-based global indexes and local indexes to support efficient similarity search and join. Since Spark is one of the widely adopted distributed in-memory computing systems, we have seamlessly integrated Dima into Spark and developed effective query optimization techniques in Spark. To the best of our knowledge, this is the first full-fledged distributed in-memory system that can support similarity-based query processing. We demonstrate our system in several scenarios, including entity matching, web table integration and query recommendation. Ji Sun 0001, Zeyuan Shang, Guoliang Li 0001, Dong Deng 0001, Zhifeng Bao |
Proc. VLDB Endow. | 2 |
| 2016 | K-Join: Knowledge-Aware Similarity JoinabstractSimilarity join is a fundamental operation in data cleaning and integration. Existing similarity-join methods utilize the string similarity to quantify the relevance but neglect the knowledge behind the data, which plays an important role in understanding the data. Thanks to public knowledge bases, e.g., Freebase and Yago, we have an opportunity to use the knowledge to improve similarity join. To address this problem, we study knowledge-aware similarity join, which, given a knowledge hierarchy and two collections of objects (e.g., documents), finds all knowledge-aware similar object pairs. To the best of our knowledge, this is the first study on knowledge-aware similarity join. There are two main challenges. The first is how to quantify the knowledge-aware similarity. The second is how to efficiently identify the similar pairs. To address these challenges, we first propose a new similarity metric to quantify the knowledge-aware similarity using the knowledge hierarchy. We then devise a filter-and-verification framework to efficiently identify the similar pairs. We propose effective signature-based filtering techniques to prune large numbers of dissimilar pairs and develop efficient verification algorithms to verify the candidates that are not pruned in the filter step. Experimental results on real-world datasets show that our method significantly outperforms baseline algorithms in terms of both efficiency and effectiveness. Zeyuan Shang, Yaxiao Liu, Guoliang Li 0001, Jianhua Feng |
IEEE Trans. Knowl. Data Eng. | 1 |