Zhengle Wang

dblp:358/8232 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
2since 2021 · last 2025
0009-0002-3493-7857ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 100%
Databases, data mining, and information retrieval
1 paper
Query processing and optimization · 100%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Performance modeling and evaluation
benchmarking
0.912025
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking · Proc. VLDB Endow. 2025
Performance modeling and evaluation
workload characterization
0.912025
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking · Proc. VLDB Endow. 2025
Performance modeling and evaluation › workload characterization
workload generation
0.912025
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking · Proc. VLDB Endow. 2025
Query processing and optimization
query workload
0.312025
PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking · Proc. VLDB Endow. 2025

Methods — techniques the papers use, named apart from their topics

timestamp assignment · 1.7multi-objective optimization · 1.7large language model · 1.7
YearPublicationVenuePosition
2025 PBench: Workload Synthesizer with Real Statistics for Cloud Analytics Benchmarking
abstract
Cloud service providers commonly use standard benchmarks like TPC-H and TPC-DS to evaluate and optimize cloud data analytics systems. However, these benchmarks rely on fixed query patterns and fail to capture real execution statistics of production cloud workloads. Although some cloud database vendors have recently released real workload traces, these traces alone do not qualify as benchmarks, as they typically lack essential components (i.e., queries and databases). To overcome this limitation, this paper studies a new problem of workload synthesis with real statistics , which generates synthetic workloads that closely approximate real execution statistics, including key performance metrics and operator distributions. To address this problem, we propose PBench, a novel workload synthesizer that constructs synthetic workloads by (1) selecting and combining workload components from existing benchmarks and (2) augmenting new workload components. This paper studies the key challenges in PBench. First, we address the challenge of balancing performance metrics and operator distributions by introducing a multi-objective optimization-based component selection method. Second, to capture the temporal dynamics of real workloads, we design a timestamp assignment method that progressively reines workload timestamps. Third, to handle the disparity between the original workload and the candidate workload, we propose a component augmentation approach that leverages large language models (LLMs) to generate additional workload components while maintaining statistical idelity. Experimental results show that PBench reduces approximation error by up to 6X compared to state-of-the-art methods.
Chunwei Liu, Bhuvan Urgaonkar, Zhengle Wang, Magnus Mueller, Chao Zhang 0034, Songyue Zhang, Pascal Pfeil, Dominik Horn, Zhengchun Liu, Davide Pagano, Tim Kraska, Samuel Madden 0001, Ju Fan
Proc. VLDB Endow.4
2024 Self-supervised Transformer-Based Pre-training Method with General Plant Infection Dataset
Zhengle Wang, Minjuan Wang, Tianyun Lai
PRCV (2)1