Tao Huang 0021

dblp:34/808-21 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2024
0009-0003-6473-3852ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2024 DashChef: A Metric Recommendation Service for Online Systems Using Graph Learning
Tao Huang 0021, Pengfei Chen 0002, Zibin Zheng
ICECCS2
2024 A Bayesian LSTM Based Active Anomaly Detection Service for Large Online Systems
abstract
Currently, many large online systems are constructed with a microservice architecture. Due to the complex dependencies, the failure of a service in such a system can cause an avalanche, which directly affects user experience and the company’s revenue. It is critical for service operators to build anomaly detection services to monitor online systems closely and comprehensively. Even though a large number of anomaly detection approaches have been proposed, few of them can simultaneously adapt to hundreds of operators’ practical detection requirements. To tackle this problem, we proposed LSTM-AAD, a Bayesian LSTM based active anomaly detection service. LSTM-AAD extracts anomaly features based on the common patterns among metrics, introduces a Bayesian LSTM model to detect anomalies in time series metrics, and employs active learning to update the online model via a small number of uncertain feedback samples. In addition, the proposed user-oriented service can be quickly responsive to operators’ further requirements. We conduct extensive experiments on real time series metrics of large online services in Tencent. The results indicate that LSTM-AAD significantly outperforms other state-of-the-art methods. Moreover, our approach can detect anomalies efficiently out of box to work in a large-scale system.
Chen Wang 0075, Tao Huang 0021, Min Li 0065, Pengfei Chen 0002
Internetware2
2024 Who is Who on Ethereum? Account Labeling Using Heterophilic Graph Convolutional Network
abstract
To combat cybercrimes and maintain financial security for the blockchain ecosystem, “know your customer” (KYC) is an essential and also challenging process due to the pseudonymity nature of blockchain technology. To unlock the potential of KYC on blockchain-based platforms like Ethereum, account labeling is a powerful means which can de-anonymize addresses by mining public transaction records. Existing studies on account labeling are mainly conducted via machine learning (ML) methods fed with hand-crafted features or graph neural networks based on the modeled transaction network. However, ML approaches based on hand-crafted features ignore the global interaction information between accounts, making it easy for criminals to evade detection. Moreover, the performance of traditional GCN methods when applied to Ethereum transaction network encounters limitations due to label sparsity, network heterophily, and large network size of the transaction network. In this article, we first analyze Ethereum accounts involved in typical businesses, in terms of both account and topological features. Then based on the analytical results, we propose a novel GCN method named know-your-customer graph convolutional network (KYC-GCN) which contains two key designs: 1) multihop aggregators and importance-based sampling are designed to tackle the dilemma between accuracy and efficiency. 2) GCN architecture is improved to explicitly capture local and more global information. Experimental results on a realistic Ethereum dataset show that the proposed KYC-GCN (90.2% accuracy, 86.2% Marco-F1) achieves state-of-the-art classification performance, and results on six benchmarks demonstrate that it yields great performance under homophily and heterophily.
Dan Lin 0007, Jiajing Wu, Tao Huang 0021, Kaixin Lin, Zibin Zheng
IEEE Trans. Syst. Man Cybern. Syst.3
2022 Share or Not Share? Towards the Practicability of Deep Models for Unsupervised Anomaly Detection in Modern Online Systems
abstract
Anomaly detection is crucial in the management of modern online systems. Due to the complexity of patterns in the monitoring data and the lack of labelled data with anomalies, recent studies mainly adopt deep unsupervised models to address this problem. Notably, even though these models have achieved a great success on experimental datasets, there are still several challenges for them to be successfully applied in a real-world modern online system. Such challenges stem from some significant properties of modern online systems, e.g., large scale, diversity and dynamics. This study investigates how these properties affect the adoption of deep anomaly detectors in modern online systems. Furthermore, we claim that model sharing is an effective way to overcome these challenges. To support this claim, we systematically study the feasibility and necessity of model sharing for unsupervised anomaly detection. In addition, we further propose a novel model, Uni-AD, which works well for model sharing. Based upon Transformer encoder layers and Base layers, Uni-AD can effectively model diverse patterns for different monitored entities and further perform anomaly detection accurately. Besides, it can accept variable-length inputs, which is a required property for a model that needs to be shared. Extensive experiments on two real-world large-scale datasets demonstrate the effectiveness and practicality of Uni-AD.
Pengfei Chen 0002, Tao Huang 0021
ISSRE3
2022 A Transferable Time Series Forecasting Service Using Deep Transformer Model for Online Systems
abstract
Many real-world online systems expect to forecast the future trend of software quality to better automate operational processes, optimize software resource cost and ensure software reliability. To achieve that, all kinds of time series metrics collected from online software systems are adopted to characterize and monitor the quality of software services. To meet relevant software engineers’ requirements, we focus on time series forecasting and aim to provide an event-driven and self-adaptive forecasting service. In this paper, we present TTSF-transformer, a transferable time series forecasting service using deep transformer model. TTSF-transformer normalizes multiple metric frequencies to ensure the model sharing across multi-source systems, employs a deep transformer model with Bayesian estimation to generate the predictive marginal distribution, and introduces transfer learning and incremental learning into the training process to ensure the performance of long-term prediction. We conduct experiments on real-world time series metrics from two different types of game business in Tencent®. The results show that TTSF-transformer significantly outperforms other state-of-the-art methods and is suitable for wide deployment in large online systems.
Tao Huang 0021, Pengfei Chen 0002, Jingrun Zhang
ASE1
2022 A Semi-Supervised VAE Based Active Anomaly Detection Framework in Multivariate Time Series for Online Systems
abstract
Nowadays, the large online systems are constructed on the basis of microservice architecture. A failure in this architecture may cause a series of failures due to the fault propagation. Thus, the large online systems need to be monitored comprehensively to ensure the service quality. Even though many anomaly detection techniques have been proposed, few of them can be directly applied to a given microservice or cloud server in industrial environment. To settle these challenges, this paper presents SLA-VAE, a semi-supervised learning based active anomaly detection framework using variational auto-encoder. SLA-VAE first defines anomalies based on feature extraction module, introduces semi-supervised VAE to identify anomalies in multivariate time series, and employs active learning to update the online model via a small number of uncertain samples. We conduct experiments on the cloud server data from two different types of game business in Tencent. The results show that SLA-VAE significantly outperforms other state-of-the-art methods and is suitable for wide deployment in large online business system.
Tao Huang 0021, Pengfei Chen 0002
WWW1