Yang Wang 0147

dblp:181/2842-147 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
5since 2021 · last 2023
0000-0002-2650-2725ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Computer networks · 8 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Grace: Interpretable Root Cause Analysis by Graph Convolutional Network for Microservices
abstract
With the development of cloud applications, large monolithic services have been replaced by loosely-coupled and single-purpose microservices. To improve localization performance, deep learning techniques have been widely used for root cause analysis. Existing research focuses on performance, yet the lack of interpretability creates key barriers to applying deep learning models in practice. This paper presents Grace, an interpretable root cause localization framework. Our work has three aims. First, to more accurately localize root causes using a Spatial-Temporal Graph Convolutional Network (STGCN). To the best of our knowledge, we are the first to apply an STGCN in the domain. Second, to design an interpreter that helps engineers to understand the system decision (beyond the simple binary black-box result). Third, we apply our interpreter to two real cases, which replaces expert knowledge to help build prior knowledge for fault type diagnosis. Our results show that Grace consistently achieves improvements over other state-of-the-art models by 4%-140%.
Yang Wang 0147, Zhenyu Li 0001, Gareth Tyson, Tianhao Miao, Gaogang Xie
IWQoS2
2023 Triple:The Interpretable Deep Learning Anomaly Detection Framework based on Trace-Metric-Log of Microservice
abstract
Existing anomaly detection approaches based on deep-learning just could simultaneously dig out key information from two dimensions in the traces, metrics or logs. Besides, they just output simple binary result, which ignores the key artificial statement information in the log. In this paper, we propose Triple, an interpretable anomaly detection approach based on deep learning for microservice system. More importantly, Triple aims to help engineers to establish trust in the system decision from key metrics and the artificial statements in logs. Triple leverages graph representation to describe the complicated dependency relationship in the traces with the logs and metrics embedded into the node features. Based on the graph representation, Triple trains a Spatial-Temporal Graph Convolutional Network(STGCN) to capture the key information and generate decision boundary by deep SVDD, which detects the system's anomaly. In addition, we design an interpreter to transfer the simple binary result into a humanly understandable result, including log, metrics and trace, to facilitate engineers' understanding and handling of the incoming incident. Our work has four aims. First, to the best of our knowledge, we are the first to simultaneously apply three data sources to finish anomaly detection in the domain. Second, we design a new anomaly detection method that is an STGCN based on SVDD. Third, we design an interpreter that makes the decision not only a simple binary result. The interpretable result could capture the key artificial statement information in the log and assist engineers in incident troubleshooting. Finally, we design a series experiments to validate our method's effectiveness in the real-world system's dataset. Our results show that Triple consistently achieves improvements over other state-of-the-art models by 11%-65%.
Yang Wang 0147, Zhenyu Li 0001, Gaogang Xie
IWQoS2
2023 AD2S: Adaptive anomaly detection on sporadic data streams
Yang Wang 0147, Zhenyu Li 0001, Hongtao Guan, Gaogang Xie
Comput. Commun.2
2023 Improving the Scalability of Distributed Network Emulations: An Algorithmic Perspective
abstract
By deploying virtualized network elements (hosts, switches, routers, links, etc.) on clusters of commodity machines, distributed network emulations (DNE) closely mimic the behaviors of network systems and provide real-time interactions and analysis for network service management. However, DNE encounters scalability challenges when faced with large network topologies. These challenges can be boiled down to the assignment problem: to which physical machine each virtualized network element should be assigned so that the largest possible network topology can be emulated? In this paper, we tackle this problem from an algorithmic perspective. We first propose TBR (topology balancing relaxation) as the relaxation of the assignment problem. TBR tries to maintain a balance of the hardware resource consumption, by minimizing the maximum inter-machine bandwidth. We further develop TBS (topology balancing solver), which combines mathematical techniques with multi-level algorithms to solve TBR efficiently. We integrate TBR and TBS into MaxiNet, a famous distributed network emulator. Experimental results show that with the same available physical resources, TBR and TBS can improve emulation scalability by up to$4.7\times $compared to baselines.
Huaiyi Zhao, Xinyi Zhang 0004, Yang Wang 0147, Zulong Diao, Yanbiao Li 0001, Gaogang Xie
IEEE Trans. Netw. Serv. Manag.3
2022 MicroCBR: Case-Based Reasoning on Spatio-temporal Fault Knowledge Graph for Microservices Troubleshooting
Yang Wang 0147, Zhenyu Li 0001, Hongtao Guan, Gaogang Xie
ICCBR2
2020 LogSayer: Log Pattern-driven Cloud Component Anomaly Diagnosis with Machine Learning
abstract
Anomaly diagnosis is a critical task for building a reliable cloud system and speeding up the system recovery form failures. With the increase of scales and applications of clouds, they are more vulnerable to various anomalies, and it is more challenging for anomaly troubleshooting. System logs that record significant events at critical time points become excellent sources of information to perform anomaly diagnosis. Never-theless, existing log-based anomaly diagnosis approaches fail to achieve high precision in highly concurrent environments due to interleaved unstructured logs. Besides, transient anomalies that have no obvious features are hard to detect by these approaches. To address this gap, this paper proposes LogSayer, a log pattern-driven anomaly detection model. LogSayer represents the system state by identifying suitable statistical features (e.g. frequency, surge), which are not sensitive to the exact log sequence. It then measures changes in the log pattern when a transient anomaly occurs. LogSayer uses Long Short-Term Memory (LSTM) neural networks to learn the historical correlation of log patterns and applies a BP neural network for adaptive anomaly decisions. Our experimental evaluations over the HDFS and OpenStack data sets show that LogSayer outperforms the state-of-the-art log-based approaches with precision over 98%.
Pengpeng Zhou, Yang Wang 0147, Zhenyu Li 0001, Xin Wang 0001, Gareth Tyson, Gaogang Xie
IWQoS2
2020 Logchain: Cloud workflow reconstruction & troubleshooting with unstructured logs
Pengpeng Zhou, Yang Wang 0147, Zhenyu Li 0001, Gareth Tyson, Hongtao Guan, Gaogang Xie
Comput. Networks2
2017 Enabling automatic composition and verification of service function chain
abstract
NFV together with SDN promises to provide more flexible and efficient service provision methods by decoupling the network functions (NFs) from the physical network topology and devices, but requires the real-time and automatic composition and verification for service function chain (SFC). However, most of SFCs today are still typically built through manual configuration processes, which are slow and error prone. In this paper, we present a novel SFC composition framework, called Automatic Composition Toolkit (ACT). It aims to automatically detect the dependencies and conflicts between NFs, so as to compose and verify SFCs before they are enforced on the physical infrastructure.
Yang Wang 0147, Zhenyu Li 0001, Gaogang Xie, Kavé Salamatian
IWQoS1
2016 Transparent flow migration for NFV
abstract
NFV together with SDN provides the flexibility for NFs in the way that they are deployed and managed. The flexibility enables dynamical scale in and scale out through migrating in-process flows among NFs. Due to stateful packet processing in NFs, flow migration has to guarantee loss-free and order-preserving for both flow states and packets. Existing frameworks closely coupled state transfer and packets migration, and thus fail to achieve safe and efficient migration with low overhead. This paper presents our design and implementation of a distributed flow migration framework, Transparent Flow Migration (TFM). TFM completely decouples the state transfer and packets migrations. The decoupling allows us to optimize the two processes separately and run them in parallel. TFM implements various optimizations through the TFM box, a shim layer providing transparent packet migration to NFs. Our evaluation shows that TFM guarantees loss-free and order-preserving for both scale-in and scale-out flow migration, and outperforms existing approaches with 3× smaller migration time. Besides, TFM uses small overhead and has very limited impacts on throughput of live TCP flows.
Yang Wang 0147, Gaogang Xie, Zhenyu Li 0001, Peng He 0003, Kavé Salamatian
ICNP1