Hao Wang 0210

dblp:181/2812-210 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2024
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Cloud and datacenter computing · 88% Performance modeling and evaluation · 12%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Cloud and datacenter computing
cluster resource management and scheduling
0.812024
Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024
Cloud and datacenter computing › configuration tuning
configuration auto-tuning
0.812024
Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024
Cloud and datacenter computing › big data platform
shuffle service
0.212024
Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024
Performance modeling and evaluation
workload characterization
0.212024
Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance · Proc. VLDB Endow. 2024

Methods — techniques the papers use, named apart from their topics

rule-based tuning · 0.8push-based shuffle · 0.8algorithm-based tuning · 0.8
YearPublicationVenuePosition
2024 Towards Resource Efficiency: Practical Insights into Large-Scale Spark Workloads at ByteDance
abstract
At ByteDance, where we execute over a million Spark jobs and handle 500PB of shuffled data daily, ensuring resource efficiency is paramount for cost savings. However, achieving optimization of resource efficiency in large-scale production environments poses significant challenges. Drawing from our practical experiences, we have identified three key issues critical to addressing resource efficiency in real-world production settings: 1 slow I/Os leading to excessive CPU and memory idleness, 2 coarse-grained resource control causing wastage, and 3 sub-optimal job configurations resulting in low utilization. To tackle these issues, we propose a resource efficiency governance framework for Spark workloads. Specifically, 1 we devise the multi-mechanism shuffle services, including Enhanced External Shuffle Service (ESS) and Cloud Shuffle Service (CSS), where CSS employs a push-based approach to enhance I/O efficiency through sequential reading. 2 We modify the Spark configuration parameter protocol, allowing for fine-grained resource control by introducing several new parameters such as milliCores and memoryBurst, as well as supporting operators with additional spill modes. 3 We design a two-stage configuration autotuning method, comprising rule-based and algorithm-based tuning, providing more reliable Spark configuration optimizations. By deploying these techniques on millions of Spark jobs in production over the last two years, we have achieved over 22% CPU utilization increase, 5% memory utilization increase, and 10% shuffle block time ratio decrease, effectively saving millions of CPU cores and petabytes of memory daily.
Xiuqi Huang, Wei Zhongjia, Hang Cheng, Chaohui Xin, Zuzhi Chen, Binbin Chen 0005, Yufei Wu 0014, Hao Wang 0210, Tieying Zhang, Xiaofeng Gao 0001, Yuming Liang, Pengwei Zhao, Guihai Chen
Proc. VLDB Endow.9