Tianhao Miao

dblp:319/0156 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2023
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer networks
1 paper
Software-defined and programmable networks · 44% Edge and fog computing · 44% Network measurement and analytics · 13%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
Performance modeling and evaluation · 50% Parallel and multicore computing · 50%

Topics — the 7 heaviest of 7, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning › distributed training › communication-efficient training › communication optimization
communication-computation overlap
0.612022
Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022
Machine learning › Efficient and distributed learning
distributed training
0.612022
Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022
Edge and fog computing › mobile edge computing
computation offloading
0.612022
Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading · IEEE Trans. Parallel Distributed Syst. 2022
Software-defined and programmable networks
programmable data plane
0.612022
Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading · IEEE Trans. Parallel Distributed Syst. 2022
Network measurement and analytics › network telemetry
in-network telemetry
0.212022
Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading · IEEE Trans. Parallel Distributed Syst. 2022
Parallel and multicore computing
distributed deep learning training
0.212022
Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022
Performance modeling and evaluation
performance prediction
0.212022
Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022

Methods — techniques the papers use, named apart from their topics

tensor fusion · 1.1recursive modeling · 1.1table approximation · 0.6logarithm projection · 0.6
YearPublicationVenuePosition
2023 Grace: Interpretable Root Cause Analysis by Graph Convolutional Network for Microservices
abstract
With the development of cloud applications, large monolithic services have been replaced by loosely-coupled and single-purpose microservices. To improve localization performance, deep learning techniques have been widely used for root cause analysis. Existing research focuses on performance, yet the lack of interpretability creates key barriers to applying deep learning models in practice. This paper presents Grace, an interpretable root cause localization framework. Our work has three aims. First, to more accurately localize root causes using a Spatial-Temporal Graph Convolutional Network (STGCN). To the best of our knowledge, we are the first to apply an STGCN in the domain. Second, to design an interpreter that helps engineers to understand the system decision (beyond the simple binary black-box result). Third, we apply our interpreter to two real cases, which replaces expert knowledge to help build prior knowledge for fault type diagnosis. Our results show that Grace consistently achieves improvements over other state-of-the-art models by 4%-140%.
Yang Wang 0147, Zhenyu Li 0001, Gareth Tyson, Tianhao Miao, Gaogang Xie
IWQoS6
2022 MD-Roofline: A Training Performance Analysis Model for Distributed Deep Learning
abstract
Due to the bulkiness and sophistication of the Distributed Deep Learning (DDL) systems, it leaves an enormous challenge for AI researchers and operation engineers to analyze, diagnose and locate the performance bottleneck during the training stage. Existing performance models and frameworks gain little insight on the performance reduction that a performance straggler induces. In this paper, we introduce MD-Roofline, a training performance analysis model, which extends the traditional rooftine model with communication dimension. The model considers the layer-wise attributes at application level, and a series of achievable peak performance metrics at hardware level. With the assistance of our MD-Roofline, the AI researchers and DDL operation engineers could locate the system bottleneck, which contains three dimensions: intra-GPU computation capacity, intra-GPU memory access bandwidth and inter-GPU communication bandwidth. We demonstrate that our performance analysis model provides great insights in bottleneck analysis when training 12 classic CNNs.
Tianhao Miao, Qinghua Wu 0004, Penglai Cui, Zhenyu Li 0001, Gaogang Xie
ISCC1
2022 Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training
abstract
Distributed Deep Learning (DDL) is widely used to accelerate deep neural network training for various Web applications. In each iteration of DDL training, each worker synchronizes neural network gradients with other workers. This introduces communication overhead and degrades the scaling performance. In this paper, we propose a recursive model, OSF (Scaling Factor considering Overlap), for estimating the scaling performance of DDL training of neural network models, given the settings of the DDL system. OSF captures two main characteristics of DDL training: the overlap between computation and communication, and the tensor fusion for batching updates. Measurements on a real-world DDL system show that OSF obtains a low estimation error (ranging from 0.5% to 8.4% for different models). Using OSF, we identify the factors that degrade the scaling performance, and propose solutions to effectively mitigate their impacts. Specifically, the proposed adaptive tensor fusion improves the scaling performance by 32.2%∼ 150% compared to the constant tensor fusion buffer size.
Tianhao Miao, Qinghua Wu 0004, Zhenyu Li 0001, Guangxin He, Jiaoren Wu, Shengzhuo Zhang, Xingwu Yang, Gareth Tyson, Gaogang Xie
WWW2
2022 Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading
abstract
Programmable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded into the network on programmable switches. These tasks may require the support of on-the-fly floating-point operations. Unfortunately, programmable switches are restricted to simple integer arithmetic operations. Existing systems circumvent this restriction by converting floats to integers or relying on local CPUs of switches, incurring extra processing delayed and accuracy loss. To address this gap, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. Specifically, NetFC utilizes logarithm projection and transformation to convert the original huge table enumerating all operands and results into several much smaller tables that can fit into the data plane of programmable switches. To cope with the table inflation problem on 32-bit floats, we also propose an approximation method that further breaks the large tables into smaller ones. In addition, NetFC leverages two optimizations to improve accuracy and reduce on-chip memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.9% with only 448KB memory consumption for 16-bit floats and 99.1% with 496KB memory consumption for 32-bit floats. Furthermore, we integrate NetFC into two distributed applications and two in-network telemetry systems to show its effectiveness in further improving the performance.
Penglai Cui, Zhenyu Li 0001, Penghao Zhang, Tianhao Miao, Jianer Zhou, Hongtao Guan, Gaogang Xie
IEEE Trans. Parallel Distributed Syst.5