VLDB 2026 Research / reviewers in the wild / expert
Wei Zhang 0172
dblp:10/4661-172
· DBLP profile ↗
5ranked-venue papers
0as first author
3since 2021 · last 2024
0000-0002-6195-169XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Resource Allocation with Service Affinity in Large-Scale Cloud EnvironmentsabstractContainerization has garnered substantial favor among cloud service providers. Nevertheless, the notable network overhead incurred between containers has prompted concerns within the community. In cloud resource scheduling, collocating service containers that frequently communicate to the same machine - termed “service affinity” - is instrumental in enhancing application performance. In response to this concern, we present a solution that harnesses service affinity and collocates containers to enhance the overall system performance and stability. To maximize the benefits of collocating containers, it is necessary to calculate a new schedule that optimally and efficiently maximizes service affinity, especially within the expansive domain of industry-scale cloud environments. In pursuit of this, we leverage the skewness property of affinity and machine learning to fuse solver-based algorithms, thereby assuring both quality and efficiency for problems at scale. Our methodology encompasses the partitioning of a given task into discrete subproblems, with a keen focus on resolving the most critical ones. Via a graph neural network classifier, we assign each subproblem to be solved independently using methods based on off-the-shelf solvers in our algorithm pool - namely, MIP-based, or column generation. This strategic approach enables the efficient computation of a schedule for a cloud cluster that fully optimizes the overall service affinity. We further propose a heuristic algorithm to compute executable container migration plans for practical use, facilitating the transition to the new placement where service affinity is well optimized. Our solution has been deployed in our large-scale production environment, covering over a million cores within ByteDance. Through the successful real-world production deployment, our approach exhibits an average improvement in end-to-end latency by 23.75% and a reduction in request error rates by 24.09% compared to the original system. Zuzhi Chen, Fuxin Jiang, Binbin Chen 0005, Yu Li 0003, Yunkai Zhang 0002, Jianjun Chen 0001, Wu Xiang, Guozhu Cheng, Wei Zhang 0172, Tieying Zhang |
ICDE | 14 |
| 2024 | ResLake: Towards Minimum Job Latency and Balanced Resource Utilization in Geo-distributed Job SchedulingabstractAt internet scale companies like ByteDance, data is generated and consumed at enormously high speed by many different applications. Achieving low latency on such big data jobs is an important problem. However, the naive approach of aggregating all the data required by a job to a single location is not always feasible in a geo-distributed environment. Similarly, existing approaches in geo-distributed job scheduling often try to minimize WAN usage, which may come at the cost of latency. Another crucial element to ensure low latency is resource load balancing among DCs, which enables flexibility in job scheduling and avoids resource bottlenecks. Therefore, to minimize latency, optimizing job completion time (JCT) while maintaining resource utilization balance is important. To this end, we propose ResLake , a global scheduling platform for data-intensive workloads. ResLake aims to reduce JCT of geo-distributed applications while balancing the compute (CPU/Memory) and storage (Disk) usages across DCs and efficiently using WAN interconnections. We have deployed ResLake in ByteDance's production for over 1.5 years. ResLake has scheduled billions of jobs since its deployment. We find that ResLake improves JCT of jobs by at least 20%, and can improve resource utilization balance across DCs by up to 53%. Xin-Chun Zhang, Aqsa Kashaf, Yihan Zou, Wei Zhang 0172, Weibo Liao, Song Haoxiang, Jintao Ye, Binbin Chen 0005, Zuzhi Chen, Tieying Zhang, Yongping Tang |
Proc. VLDB Endow. | 4 |
| 2023 | A non-convex piecewise quadratic approximation of ℓ 0 regularization: theory and accelerated algorithm
Wei Zhang 0172, Yanqin Bai |
J. Glob. Optim. | 2 |
| 2018 | A P-ADMM for sparse quadratic kernel-free least squares semi-supervised support vector machine
Yaru Zhan, Yanqin Bai, Wei Zhang 0172, Shihui Ying |
Neurocomputing | 3 |
| 2017 | The Spectral Radius and Domination Number of Uniform Hypergraphs
Liying Kang, Wei Zhang 0172, Erfang Shan |
COCOA (2) | 2 |