Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Shichao Liu 0004

dblp:134/5661-4 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2023
0000-0003-4714-3749ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 61% Performance modeling and evaluation · 30% Distributed systems · 9%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
High-performance computing › system monitoring
i/o monitoring
0.712023
End-to-end I/O Monitoring on Leading Supercomputers · ACM Trans. Storage 2023
Performance modeling and evaluation
performance monitoring
0.712023
End-to-end I/O Monitoring on Leading Supercomputers · ACM Trans. Storage 2023
High-performance computing
supercomputing
0.712023
End-to-end I/O Monitoring on Leading Supercomputers · ACM Trans. Storage 2023
Distributed systems › network management
network monitoring
0.212023
End-to-end I/O Monitoring on Leading Supercomputers · ACM Trans. Storage 2023

Methods — techniques the papers use, named apart from their topics

trace compression · 0.7distributed caching · 0.7
YearPublicationVenuePosition
2023 End-to-end I/O Monitoring on Leading Supercomputers
abstract
This paper offers a solution to overcome the complexities of production system I/O performance monitoring. We present Beacon, an end-to-end I/O resource monitoring and diagnosis system for the 40960-node Sunway TaihuLight supercomputer, currently the fourth-ranked supercomputer in the world. Beacon simultaneously collects and correlates I/O tracing/profiling data from all the compute nodes, forwarding nodes, storage nodes, and metadata servers. With mechanisms such as aggressive online and offline trace compression and distributed caching/storage, it delivers scalable, low-overhead, and sustainable I/O diagnosis under production use. With Beacon’s deployment on TaihuLight for more than three years, we demonstrate Beacon’s effectiveness with real-world use cases for I/O performance issue identification and diagnosis. It has already successfully helped center administrators identify obscure design or configuration flaws, system anomaly occurrences, I/O performance interference, and resource under- or over-provisioning problems. Several of the exposed problems have already been fixed, with others being currently addressed. Encouraged by Beacon’s success in I/O monitoring, we extend it to monitor interconnection networks, which is another contention point on supercomputers. In addition, we demonstrate Beacon’s generality by extending it to other supercomputers. Both Beacon codes and part of collected monitoring data are released. 1
Bin Yang 0043, Wei Xue 0003, Shichao Liu 0004, Xiaosong Ma, Xiyang Wang 0003
ACM Trans. Storage4