Shangming Cai

dblp:263/3707 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
2since 2021 · last 2025
0000-0002-0902-7774ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Software engineering, system software, and programming languages
1 paper
Software maintenance and evolution · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Distributed systems · 67% Storage systems · 33%
Computer networks
1 paper
Wireless networking · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Software maintenance and evolution
AIOps
0.912025
LogLabeler: Towards Effective Acquisition of Log Labels in Industrial Log-Based Analysis · IEEE Trans. Serv. Comput. 2025
Software maintenance and evolution
log analysis
0.912025
LogLabeler: Towards Effective Acquisition of Log Labels in Industrial Log-Based Analysis · IEEE Trans. Serv. Comput. 2025
Wireless networking › medium access control
communication scheduling
0.612022
DynaComm: Accelerating Distributed CNN Training Between Edges and Clouds Through Dynamic Communication Scheduling · IEEE J. Sel. Areas Commun. 2022
Distributed systems
consensus
0.412020
CRaft: An Erasure-coding-supported Version of Raft for Reducing Storage Cost and Network Cost · FAST 2020
Storage systems › storage reliability
erasure coding
0.412020
CRaft: An Erasure-coding-supported Version of Raft for Reducing Storage Cost and Network Cost · FAST 2020
Distributed systems › consensus › leader-based consensus
raft
0.412020
CRaft: An Erasure-coding-supported Version of Raft for Reducing Storage Cost and Network Cost · FAST 2020
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
parameter server
0.212022
DynaComm: Accelerating Distributed CNN Training Between Edges and Clouds Through Dynamic Communication Scheduling · IEEE J. Sel. Areas Commun. 2022

Methods — techniques the papers use, named apart from their topics

dynamic scheduling · 1.1human-in-the-loop refinement · 0.9automated label generation · 0.9
YearPublicationVenuePosition
2025 LogLabeler: Towards Effective Acquisition of Log Labels in Industrial Log-Based Analysis
abstract
Log-based AIOps is a widely researched topic aiming at reducing the developer burden in system maintenance. Since industrial developers prefer lightweight supervised solutions for log-based AIOps, the strong dependence of these solutions on labeled data creates significant challenges for teams new to building log-based AIOps capabilities, such as high labeling costs, inconsistent annotations, and manual management issues. Log-labeling faces challenges in integrating existing artifacts to reduce labeling costs and manage labels effectively. To the best of our knowledge, no prior research addresses of assisting log-labeling problem. In this article, we propose a new approach called LogLabeler to assist developers in annotating and managing log labels. LogLabeler leverages existing artifacts for initial-label-acquisition, minimizes labeling costs by automatically generating all log labels, and shields developers from manual label management through a human-in-the-loop refinement approach. Evaluations on real-world datasets from Alibaba and open-source datasets show that LogLabeler can effectively supplement log labels, achieving comparable accuracy to existing baselines while operating more efficiently. Furthermore, we demonstrate LogLabeler's practical effectiveness at Alibaba through a case study, highlighting its benefits to developers.
Zongyang Li, Qinglong Wang 0003, Shangming Cai, Zheng Liu 0022, Tao Ma 0006, Wei Yang 0013, Ying Li 0012, Tao Xie 0001
IEEE Trans. Serv. Comput.3
2022 DynaComm: Accelerating Distributed CNN Training Between Edges and Clouds Through Dynamic Communication Scheduling
abstract
To reduce uploading bandwidth and address privacy concerns, deep learning at the network edge has been an emerging topic. Typically, edge devices collaboratively train a shared model using real-time generated data through the Parameter Server framework. Although all the edge devices can share the computing workloads, the distributed training processes over edge networks are still time-consuming due to the parameters and gradients transmission procedures between parameter servers and edge devices. Focusing on accelerating distributed Convolutional Neural Networks (CNNs) training at the network edge, we present DynaComm, a novel scheduler that dynamically decomposes each transmission procedure into several segments to achieve optimal layer-wise communications and computations overlapping during run-time. Through experiments, we verify that DynaComm manages to achieve optimal layer-wise scheduling for all cases compared to competing strategies while the model accuracy remains untouched.
Shangming Cai, Dongsheng Wang 0002, Haixia Wang 0001, Yongqiang Lyu 0001, Guangquan Xu, James Xi Zheng, Athanasios V. Vasilakos
IEEE J. Sel. Areas Commun.1
2020 CRaft: An Erasure-coding-supported Version of Raft for Reducing Storage Cost and Network Cost
Zizhong Wang, Tongliang Li, Haixia Wang 0001, Airan Shao, Yunren Bai, Shangming Cai, Dongsheng Wang 0002
FAST6
2020 CARD: A Congestion-Aware Request Dispatching Scheme for Replicated Metadata Server Cluster
abstract
Replicated metadata server cluster (RMSC) is highly efficient to be used in distributed filesystems while facing data-driven scenarios (e.g., massive-scale distributed machine learning tasks). Yet, when considering cost-effectiveness and system utilization, the cluster scale is commonly restricted in practice. Within this context, servers in the cluster start to suffer from load-oscillations at higher system utilization due to clients’ congestion-unaware behaviors and unintelligent selection strategies (i.e., servers in the cluster are preferred then evaded intermittently). The consequences brought by load-oscillations degrade the overall performance of the whole system to some extent. One solution to tackle this problem is having clients share a part of the responsibility and behave more wisely for stability concerns. So in this paper, we present a Congestion-Aware Request Dispatching scheme, CARD, which is mainly conducted at clients and directed by a rate control mechanism. Through extensive experiments, we verify that CARD is highly efficient in resolving load-oscillations in RMSC. Apart from this, our results show that RMSC with our congestion-aware based optimization achieves better scalability compared to previous implementations under targeted workloads, especially in heterogeneous environments.
Shangming Cai, Dongsheng Wang 0002, Zhanye Wang, Haixia Wang 0001
ICPP1