Yaocheng Xiang

dblp:217/6693 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
1since 2021 · last 2023
0000-0003-4664-3979ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
Memory systems · 70% Processor architecture and microarchitecture · 23% Performance modeling and evaluation · 7%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Memory systems
cache management
0.312018
DCAPS: dynamic cache allocation with partial sharing · EuroSys 2018
Memory systems › cache management
cache partitioning
0.312018
DCAPS: dynamic cache allocation with partial sharing · EuroSys 2018
Processor architecture and microarchitecture
multicore design
0.312018
DCAPS: dynamic cache allocation with partial sharing · EuroSys 2018
Memory systems › cache management
shared cache management
0.312018
DCAPS: dynamic cache allocation with partial sharing · EuroSys 2018
Performance modeling and evaluation
workload characterization
0.112018
DCAPS: dynamic cache allocation with partial sharing · EuroSys 2018

Methods — techniques the papers use, named apart from their topics

simulated annealing · 0.3online miss rate curve prediction · 0.3
YearPublicationVenuePosition
2023 FLORIA: A Fast and Featherlight Approach for Predicting Cache Performance
abstract
The cache Miss Ratio Curve (MRC) serves a variety of purposes such as cache partitioning, application profiling and code tuning. In this work, we propose a new metric, called cache miss distribution, that describes cache miss behavior over cache sets, for predicting cache MRCs. Based on this metric, we present FLORIA, a software-based, online approach that approximates cache MRCs on commodity systems. By polluting a tunable number of cache lines in some selected cache sets using our designed microbenchmark, the cache miss distribution for the target workload is obtained via hardware performance counters with the support of precise event based sampling (PEBS). A model is developed to predict the MRC of the target workload based on its cache miss distribution.
Jun Xiao 0009, Yaocheng Xiang, Xiaolin Wang 0001, Yingwei Luo, Andy D. Pimentel, Zhenlin Wang 0003
ICS2
2019 EMBA: Efficient Memory Bandwidth Allocation to Improve Performance on Intel Commodity Processor
abstract
On multi-core processors, contention on shared resources such as the last level cache (LLC) and memory bandwidth may cause serious performance degradation, which makes efficient resource allocation a critical issue in data centers. Intel recently introduces Memory Bandwidth Allocation (MBA) technology on its Xeon scalable processors, which makes it possible to allocate memory bandwidth in a real system. However, how to make the most of MBA to improve system performance remains an open question. In this work, (1) we formulate a quantitative relationship between a program's performance and its LLC occupancy and memory request rate on commodity processors. (2) Guided by the performance formula, we propose a heuristic bound-aware throttling algorithm to improve system performance and (3) we further develop a hierarchical clustering method to improve the algorithm's efficiency. (4) We implement these algorithms in EMBA, a low-overhead dynamic memory bandwidth scheduling system to improve performance on Intel commodity processors. The results show that, when multiple programs run simultaneously on a multi-core processor whose memory bandwidth is saturated, the programs with high memory bandwidth demand usually use bandwidth inefficiently compared with programs with medium memory bandwidth demand from the perspective of CPU performance. By slightly throttling the former's bandwidth, we can significantly improve the performance of the latter. On average, we improve system performance by 36.9% at the expense of 8.6% bandwidth utilization rate.
Yaocheng Xiang, Chencheng Ye 0001, Xiaolin Wang 0001, Yingwei Luo, Zhenlin Wang 0003
ICPP1
2018 DCAPS: dynamic cache allocation with partial sharing
abstract
In a multicore system, effective management of shared last level cache (LLC), such as hardware/software cache partitioning, has attracted significant research attention. Some eminent progress is that Intel introduced Cache Allocation Technology (CAT) to its commodity processors recently. CAT implements way partitioning and provides software interface to control cache allocation. Unfortunately, CAT can only allocate at way level, which does not scale well for a large thread or program count to serve their various performance goals effectively. This paper proposes Dynamic Cache Allocation with Partial Sharing (DCAPS), a framework that dynamically monitors and predicts a multi-programmed workload's cache demand, and reallocates LLC given a performance target. Further, DCAPS explores partial sharing of a cache partition among programs and thus practically achieves cache allocation at a finer granularity. DCAPS consists of three parts: (1) Online Practical Miss Rate Curve (OPMRC), a low-overhead software technique to predict online miss rate curves (MRCs) of individual programs of a workload; (2) a prediction model that estimates the LLC occupancy of each individual program under any CAT allocation scheme; (3) a simulated annealing algorithm that searches for a near-optimal CAT scheme given a specific performance goal. Our experimental results show that DCAPS is able to optimize for a wide range of performance targets and can scale to a large core count.
Yaocheng Xiang, Xiaolin Wang 0001, Zihui Huang, Yingwei Luo, Zhenlin Wang 0003
EuroSys1