Shanpei Chen

dblp:284/4873 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
2since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
2 papers
Cloud and datacenter computing · 71% Memory systems · 29%
Software engineering, system software, and programming languages
1 paper
Operating systems · 100%

Topics — the 4 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Operating systems › resource management › process management › CPU scheduling
kernel scheduling
0.712023
Efficient Scheduler Live Update for Linux Kernel with Modularization · ASPLOS (3) 2023
Cloud and datacenter computing
cluster resource management and scheduling
0.412020
Alita: comprehensive performance isolation through bias resource management for public clouds · SC 2020
Memory systems
memory bandwidth
0.412020
Alita: comprehensive performance isolation through bias resource management for public clouds · SC 2020
Cloud and datacenter computing
performance isolation
0.412020
Alita: comprehensive performance isolation through bias resource management for public clouds · SC 2020

Methods — techniques the papers use, named apart from their topics

stack inspection · 1.3data rebuild · 1.3hardware performance counter monitoring · 0.4adaptive throttling · 0.4
YearPublicationVenuePosition
2023 Efficient Scheduler Live Update for Linux Kernel with Modularization
abstract
The scheduler is a critical component of the operating system (OS)and is tightly coupled with Linux. Production-level clouds often host various workloads, and these workloads require different schedulers to achieve high performance. Thus the capability of updating the scheduler lively without rebooting the OS is crucial for the production environments. However, emerging live update techniques only apply for the fine-grained function-level updates or require extra constraints such as microkernel. It fails to update the entire heavy process scheduler subsystem lively. We therefore propose Plugsched to enable scheduler live update, and there are two key novelties. First of all, with the idea of modularization, Plugsched decouples the scheduler from the Linux kernel to be an independent module; Secondly, Plugsched uses the data rebuild technique to migrate the state from the old scheduler to the new one. This scheme can be directly applied to the Linux kernel scheduler in production environments without modifying kernel code. Unlike current function-level live update solutions, Plugsched allows developers to update the entire scheduler subsystem and modify internal scheduler data via the rebuilding technique. Moreover, an optimized stack inspection method is introduced to further effectively reduce the downtime due to the update. Experimental and production results show that Plugsched can effectively update kernel scheduler lively and the downtime is less than tens of milliseconds.
Teng Ma 0006, Shanpei Chen, Erwei Deng, Quan Chen 0002, Minyi Guo
ASPLOS (3)2
2023 Kronos: towards bus contention-aware job scheduling in warehouse scale computers
Shang Zhao 0003, Quan Chen 0002, Shanpei Chen, Tao Ma 0006, Yong Yang 0013, Wenli Zheng, Minyi Guo
Frontiers Comput. Sci.5
2020 Alita: comprehensive performance isolation through bias resource management for public clouds
abstract
The tenants of public cloud platforms share hard-ware resources on the same node, resulting in the potential for performance interference (or malicious attacks). A tenant is able to degrade the performance of its neighbors on the same node significantly through overuse of the shared memory bus, last level cache (LLC)/memory bandwidth, and power. To eliminate such unfairness we propose Alita, a runtime system consisting of an online interference identifier and adaptive interference eliminator. The interference identifier monitors hardware and system-level event statistics to identify resource polluters. The eliminator improves the performance of normal applications by throttling only the resource usage of polluters. Specifically, Alita adopts bus lock sparsification, bias LLC/bandwidth isolation, and selective power throttling to throttle the resource usage of polluters. Results for an experimental platform and in-production cloud platform with 30,000 nodes demonstrate that Alita significantly improves the performance of co-located virtual machines in the presence of resource polluters based on system-level knowledge.
Quan Chen 0002, Shang Zhao 0003, Shanpei Chen, Tao Ma 0006, Yong Yang 0013, Minyi Guo
SC4