EDBT 2026 Demo / reviewers in the wild / expert
Tianhao Miao
dblp:319/0156
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer networks
1 paper |
Software-defined and programmable networks · 44% Edge and fog computing · 44% Network measurement and analytics · 13% | |
| Artificial intelligence
1 paper |
Efficient and distributed learning · 100% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
Performance modeling and evaluation · 50% Parallel and multicore computing · 50% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Efficient and distributed learning › distributed training › communication-efficient training › communication optimization
communication-computation overlap |
0.6 | 1 | 2022 | Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022 |
Machine learning › Efficient and distributed learning
distributed training |
0.6 | 1 | 2022 | Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022 |
Edge and fog computing › mobile edge computing
computation offloading |
0.6 | 1 | 2022 | Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading · IEEE Trans. Parallel Distributed Syst. 2022 |
Software-defined and programmable networks
programmable data plane |
0.6 | 1 | 2022 | Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading · IEEE Trans. Parallel Distributed Syst. 2022 |
Network measurement and analytics › network telemetry
in-network telemetry |
0.2 | 1 | 2022 | Enabling In-Network Floating-Point Arithmetic for Efficient Computation Offloading · IEEE Trans. Parallel Distributed Syst. 2022 |
Parallel and multicore computing
distributed deep learning training |
0.2 | 1 | 2022 | Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022 |
Performance modeling and evaluation
performance prediction |
0.2 | 1 | 2022 | Modeling and Optimizing the Scaling Performance in Distributed Deep Learning Training · WWW 2022 |
Methods — techniques the papers use, named apart from their topics
tensor fusion · 1.1recursive modeling · 1.1table approximation · 0.6logarithm projection · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Grace: Interpretable Root Cause Analysis by Graph Convolutional Network for MicroservicesabstractWith the development of cloud applications, large monolithic services have been replaced by loosely-coupled and single-purpose microservices. To improve localization performance, deep learning techniques have been widely used for root cause analysis. Existing research focuses on performance, yet the lack of interpretability creates key barriers to applying deep learning models in practice. This paper presents Grace, an interpretable root cause localization framework. Our work has three aims. First, to more accurately localize root causes using a Spatial-Temporal Graph Convolutional Network (STGCN). To the best of our knowledge, we are the first to apply an STGCN in the domain. Second, to design an interpreter that helps engineers to understand the system decision (beyond the simple binary black-box result). Third, we apply our interpreter to two real cases, which replaces expert knowledge to help build prior knowledge for fault type diagnosis. Our results show that Grace consistently achieves improvements over other state-of-the-art models by 4%-140%. Yang Wang 0147, Zhenyu Li 0001, Gareth Tyson, Tianhao Miao, Gaogang Xie |
IWQoS | 6 |
| 2022 | MD-Roofline: A Training Performance Analysis Model for Distributed Deep LearningabstractDue to the bulkiness and sophistication of the Distributed Deep Learning (DDL) systems, it leaves an enormous challenge for AI researchers and operation engineers to analyze, diagnose and locate the performance bottleneck during the training stage. Existing performance models and frameworks gain little insight on the performance reduction that a performance straggler induces. In this paper, we introduce MD-Roofline, a training performance analysis model, which extends the traditional rooftine model with communication dimension. The model considers the layer-wise attributes at application level, and a series of achievable peak performance metrics at hardware level. With the assistance of our MD-Roofline, the AI researchers and DDL operation engineers could locate the system bottleneck, which contains three dimensions: intra-GPU computation capacity, intra-GPU memory access bandwidth and inter-GPU communication bandwidth. We demonstrate that our performance analysis model provides great insights in bottleneck analysis when training 12 classic CNNs. Tianhao Miao, Qinghua Wu 0004, Penglai Cui, Zhenyu Li 0001, Gaogang Xie |
ISCC | 1 |
| 2022 | Modeling and Optimizing the Scaling Performance in Distributed Deep Learning TrainingabstractDistributed Deep Learning (DDL) is widely used to accelerate deep neural network training for various Web applications. In each iteration of DDL training, each worker synchronizes neural network gradients with other workers. This introduces communication overhead and degrades the scaling performance. In this paper, we propose a recursive model, OSF (Scaling Factor considering Overlap), for estimating the scaling performance of DDL training of neural network models, given the settings of the DDL system. OSF captures two main characteristics of DDL training: the overlap between computation and communication, and the tensor fusion for batching updates. Measurements on a real-world DDL system show that OSF obtains a low estimation error (ranging from 0.5% to 8.4% for different models). Using OSF, we identify the factors that degrade the scaling performance, and propose solutions to effectively mitigate their impacts. Specifically, the proposed adaptive tensor fusion improves the scaling performance by 32.2%∼ 150% compared to the constant tensor fusion buffer size. Tianhao Miao, Qinghua Wu 0004, Zhenyu Li 0001, Guangxin He, Jiaoren Wu, Shengzhuo Zhang, Xingwu Yang, Gareth Tyson, Gaogang Xie |
WWW | 2 |
| 2022 | Enabling In-Network Floating-Point Arithmetic for Efficient Computation OffloadingabstractProgrammable switches are recently used for accelerating data-intensive distributed applications. Some computational tasks, traditionally performed on servers in data centers, are offloaded into the network on programmable switches. These tasks may require the support of on-the-fly floating-point operations. Unfortunately, programmable switches are restricted to simple integer arithmetic operations. Existing systems circumvent this restriction by converting floats to integers or relying on local CPUs of switches, incurring extra processing delayed and accuracy loss. To address this gap, we propose NetFC, a table-lookup method to achieve on-the-fly in-network floating-point arithmetic operations nearly without accuracy loss. Specifically, NetFC utilizes logarithm projection and transformation to convert the original huge table enumerating all operands and results into several much smaller tables that can fit into the data plane of programmable switches. To cope with the table inflation problem on 32-bit floats, we also propose an approximation method that further breaks the large tables into smaller ones. In addition, NetFC leverages two optimizations to improve accuracy and reduce on-chip memory consumption. We use both synthetic and real-life datasets to evaluate NetFC. The experimental results show that the average accuracy of NetFC is above 99.9% with only 448KB memory consumption for 16-bit floats and 99.1% with 496KB memory consumption for 32-bit floats. Furthermore, we integrate NetFC into two distributed applications and two in-network telemetry systems to show its effectiveness in further improving the performance. Penglai Cui, Zhenyu Li 0001, Penghao Zhang, Tianhao Miao, Jianer Zhou, Hongtao Guan, Gaogang Xie |
IEEE Trans. Parallel Distributed Syst. | 5 |