Haoyu Liu 0002

dblp:203/9503-2 · DBLP profile ↗
← Back
24ranked-venue papers
3as first author
20since 2021 · last 2025
0000-0002-8998-1217ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Computer networks · 6 · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021
YearPublicationVenuePosition
2025 MAIN: Mutual Alignment Is Necessary for instruction tuning
abstract
Fanyi Yang, Jianfeng Liu, Xin Zhang, Haoyu Liu, Xixin Cao, Yuefeng Zhan, Hao Sun, Weiwei Deng, Feng Sun, Qi Zhang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Fanyi Yang, Xin Zhang 0099, Haoyu Liu 0002, Xixin Cao, Yuefeng Zhan, Hao Sun 0015, Feng Sun 0008, Qi Zhang 0066
EMNLP4
2025 Empowering Economic Simulation for Massively Multiplayer Online Games through Generative Agent-Based Modeling
abstract
Within the domain of Massively Multiplayer Online (MMO) economy research, Agent-Based Modeling (ABM) has emerged as a robust tool for analyzing game economics, evolving from rule-based agents to decision-making agents enhanced by reinforcement learning. Nevertheless, existing works encounter significant challenges when attempting to emulate human-like economic activities among agents, particularly regarding agent reliability, sociability, and interpretability.In this study, we take a preliminary step in introducing a novel approach using Large Language Models (LLMs) in MMO economy simulation. Leveraging LLMs' role-playing proficiency, generative capacity, and reasoning aptitude, we design LLM-driven agents with human-like decision-making and adaptability. These agents are equipped with the abilities of role-playing, perception, memory, and reasoning, addressing the aforementioned challenges effectively. Simulation experiments focusing on in-game economic activities demonstrate that LLM-empowered agents can promote emergent phenomena like role specialization and price fluctuations in line with market rules.
Bihan Xu, Runze Wu 0001, Zhenya Huang, Zhipeng Hu, Kai Wang 0064, Haoyu Liu 0002, Tangjie Lv, Changjie Fan, Xin T. Tong, Jiangze Han
KDD (2)8
2025 Intention-Aware Denoising Diffusion Model for Trajectory Prediction
abstract
Trajectory prediction is an essential component in autonomous driving, particularly for collision avoidance systems. Considering the inherent uncertainty of the task, numerous studies have utilized generative models to produce multiple plausible future trajectories for each agent. However, most of them suffer from limited representation ability or unstable training issues. To overcome these limitations, we propose utilizing the diffusion model to generate the distribution of future trajectories. Two cruxes are to be settled to realize such an idea. First, the diversity of intention is intertwined with the uncertain surroundings, making the true distribution hard to parameterize. Second, the diffusion process is time-consuming during the inference phase, rendering it unrealistic to implement in a real-time driving system. We propose an Intention-aware denoising Diffusion Model (IDM), which addresses the above two problems. We decouple the original uncertainty into intention uncertainty and action uncertainty and model them with two dependent diffusion processes. To decrease the inference time, we reduce the variable dimensions in the intention-aware diffusion process and restrict the initial distribution of the action-aware diffusion process, which leads to fewer diffusion steps. To validate our approach, we conduct experiments on the Stanford Drone Dataset (SDD) and the ETH/UCY dataset. Our methods achieve state-of-the-art results, with a minFDE of 13.83 pixels on the SDD dataset and 0.36 meters on ETH/UCY datasets. Compared with the original diffusion model, IDM reduces inference time by two-thirds. Interestingly, our experiments further reveal that introducing intention information is beneficial in modeling the diffusion process of fewer steps.
Chen Liu 0034, Shibo He, Haoyu Liu 0002, Jiming Chen 0001
IEEE Trans. Intell. Transp. Syst.3
2024 EnMatch: Matchmaking for Better Player Engagement via Neural Combinatorial Optimization
abstract
Matchmaking is a core task in e-sports and online games, as it contributes to player engagement and further influences the game's lifecycle. Previous methods focus on creating fair games at all times. They divide players into different tiers based on skill levels and only select players from the same tier for each game. Though this strategy can ensure fair matchmaking, it is not always good for player engagement. In this paper, we propose a novel Engagement-oriented Matchmaking (EnMatch) framework to ensure fair games and simultaneously enhance player engagement. Two main issues need to be addressed. First, it is unclear how to measure the impact of different team compositions and confrontations on player engagement during the game considering the variety of player characteristics. Second, such a detailed consideration on every single player during matchmaking will result in an NP-hard combinatorial optimization problem with non-linear objectives. In light of these challenges, we turn to real-world data analysis to reveal engagement-related factors. The resulting insights guide the development of engagement modeling, enabling the estimation of quantified engagement before a match is completed. To handle the combinatorial optimization problem, we formulate the problem into a reinforcement learning framework, in which a neural combinatorial optimization problem is built and solved. The performance of EnMatch is finally demonstrated through the comparison with other state-of-the-art methods based on several real-world datasets and online deployments on two games.
Kai Wang 0064, Haoyu Liu 0002, Zhipeng Hu, Xiaochuan Feng, Minghao Zhao 0002, Runze Wu 0001, Tangjie Lv, Changjie Fan
AAAI2
2024 Attention-based Vision Knowledge Adaptation for Constrained Continual Learning
abstract
The demand for continual machine learning in the context of limited computational resources and data availability is critical in the evolving landscape of the connected digital world. Current network applications predominantly rely on deep learning models that require labor/computation-intensive training processes. These models often struggle to effectively adapt to new data while preserving performance on previously acquired knowledge. In this paper, we introduce a lightweight framework for continual knowledge adaptation and learning designed to address these challenges. To prevent disruption of existing services, we propose an attention-based adapter that integrates seamlessly with the existing vision model to encode new incoming data. The weights of the original model are kept fixed during the adaptation process, ensuring the preservation of previously learned knowledge. Furthermore, to enhance learning efficiency and accelerate convergence with new data, we implement a knowledge fusion mechanism that facilitates interaction between existing knowledge and information from new data. Our framework is modular, enabling flexible deployment across distributed devices. The adapter and knowledge fusion module are implemented at each stage with minimal trainable parameters, optimizing resource usage. Extensive experiments and ablation studies validate the effectiveness of the proposed framework.
Bicheng Guo, Conghao Zhou, Haoyu Liu 0002, Shibo He, Jiming Chen 0001, Xuemin Shen
GLOBECOM3
2024 Treemil: A Multi-Instance Learning Framework for Time Series Anomaly Detection with Inexact Supervision
abstract
Time series anomaly detection (TSAD) plays a vital role in various domains such as healthcare, networks and industry. Considering labels are crucial for detection but difficult to obtain, we turn to TSAD with inexact supervision: only series-level labels are provided during the training phase, while point-level anomalies are predicted during the testing phase. Previous works follow a traditional multi-instance learning (MIL) approach, which focuses on encouraging high anomaly scores at individual time steps. However, time series anomalies are not only limited to individual point anomalies, they can also be collective anomalies, typically exhibiting abnormal patterns over subsequences. To address the challenge of collective anomalies, in this paper, we propose a tree-based MIL framework (TreeMIL). We first adopt an N-ary tree structure to divide the entire series into multiple nodes, where nodes at different levels represent subsequences with different lengths. Then, the subsequences’ features are extracted to determine the presence of collective anomalies. Finally, we calculate point-level anomaly scores by aggregating features from nodes at different levels. Experiments conducted on seven public datasets and eight baselines demonstrate that TreeMIL achieves an average 32.3% improvement in F1-score compared to previous state-of-the-art methods. The code is available at https://github.com/fly-orange/TreeMIL.
Chen Liu 0034, Shibo He, Haoyu Liu 0002, Shizhong Li
ICASSP3
2024 XRL-Bench: A Benchmark for Evaluating and Comparing Explainable Reinforcement Learning Techniques
abstract
Reinforcement Learning (RL) has demonstrated substantial potential across diverse fields, yet understanding its decision-making process, especially in real-world scenarios where rationality and safety are paramount, is an ongoing challenge. This paper delves in to Explainable RL (XRL), a subfield of Explainable AI (XAI) aimed at unravelling the complexities of RL models. Our focus rests on state-explaining techniques, a crucial subset within XRL methods, as they reveal the underlying factors influencing an agent's actions at any given time. Despite their significant role, the lack of a unified evaluation framework hinders assessment of their accuracy and effectiveness. To address this, we introduce XRL-Bench, a unified standardized benchmark tailored for the evaluation and comparison of XRL methods, encompassing three main modules: standard RL environments, explainers based on state importance, and standard evaluators. XRL-Bench supports both tabular and image data for state explanation. We also propose TabularSHAP, an innovative and competitive XRL method. We demonstrate the practical utility of TabularSHAP in real-world online gaming services and offer an open-source benchmark platform for the straightforward implementation and evaluation of XRL methods. Our contributions facilitate the continued progression of XRL technology.
Zhipeng Hu, Runze Wu 0001, Xingchen Fang, Ji Jiang, Tianze Zhou, Yujing Hu, Haoyu Liu 0002, Tangjie Lyu, Changjie Fan
KDD10
2024 Toward Efficient Traffic Incident Detection via Explicit Edge-Level Incident Modeling
abstract
Traffic incident detection is a critical task within traffic monitoring systems, enabling on-the-fly alerts for emergency actions. Numerous efforts have been made to detect and localize traffic incidents using data recorded by inductive loop detectors. However, they only focus on the node-level incidents that happen within the surveillance areas and ignore the edge-level ones that take place outside of these areas. In this paper, we propose to detect both kinds of incidents simultaneously based on the sparsely distributed sensors. An important challenge is how to explicitly model the edge status and detect this kind of incidents. Additionally, capturing complex relationships among traffic dynamics, road locations, and temporal information is non-trivial. In this paper, we first describe the traffic dynamics by a fine-grained graph where the sensor range is designed as a hyper-parameter to control the coverage boundaries. Then, we propose an Edge-and Node-aware Dual AutoEncoder (ENDAE), where the correlations are decoupled into inter-nodes, inter-series and inter-attribute parts, which are further captured via node encoder, temporal encoder and attribute encoder, respectively. Furthermore, the reconstruction errors are calculated for node-level and edge-level event detection separately. The overall method is evaluated based on two real-world datasets from Bay Area and Los Angeles in California. ENDAE surpasses all the state-of-the-art method in both kinds of incidents, with at least 12.5% improvement in recall and 18.5% decrease in delay. Notably, for edge-level incidents, ENDAE achieves double the recall of the previous SOTA methods.
Chen Liu 0034, Jiming Chen 0001, Haoyu Liu 0002, Shizhong Li, Shibo He
IEEE Internet Things J.3
2024 Promoting human-AI interaction makes a better adoption of deep reinforcement learning: a real-world application in game industry
Zhipeng Hu, Haoyu Liu 0002, Lizi Wang, Runze Wu 0001, Yujing Hu, Tangjie Lyu, Changjie Fan
Multim. Tools Appl.2
2024 Latency-Aware Neural Architecture Performance Predictor With Query-to-Tier Technique
abstract
Neural Architecture Search (NAS) is a powerful tool for automating effective image and video processing DNN designing. The ranking of the accuracy has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between the two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. On the contrary, we propose to let the performance predictor concentrate on the global quality level of specific architecture, and learn the tier embeddings of the whole search space automatically with learnable queries. The proposed method, dubbed as Neural Architecture Ranker with Query-to-Tier technique (NARQ2T), explores the quality tiers of the search space globally and classifies each individual to the tier they belong to. Thus, the predictor gains knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Thanks to the encoder-decoder design, our method is able to predict the latency of the searched model without deteriorating the performance prediction. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning or Evolutionary Algorithm, thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NARQ2T achieves state-of-the-art performance on two widely used datasets for NAS research. Moreover, extensive experiments have validated the efficacy of the designed method.
Bicheng Guo, Lilin Xu, Tao Chen 0003, Peng Ye 0006, Shibo He, Haoyu Liu 0002, Jiming Chen 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 WindTrans: Transformer-Based Wind Speed Forecasting Method for High-Speed Railway
abstract
Wind speed forecasting provides the upcoming wind information and is important to the safe operation of High-Speed Railway (HSR). However, it remains a challenge due to the stochastic and highly varying characteristics of wind. In this paper, we propose a novel Transformer-based method for short-term wind speed forecasting, named WindTrans. Two major cruxes are addressed. First, the task is performed on fine-grained wind speed gathered from multiple sensors. These data present dynamic intra-series and inter-series correlations, which are hard for previous methods to recover. We advance a Transformer-based deep learning model, which has two distinctive characteristics: (1) a graph encoder, which captures the dynamic spatial correlation among wind speeds at different locations, and (2) a temporal decoder to model long sequence wind speed time series, which is resistant to noise in time series. Second, wind speed patterns gradually evolve in long-term periods, thus deactivating prediction models trained on historical data. To tackle this bottleneck, we put forward an experience replay-based scheme to renew the model regularly. To ensure that the renewed model still dominates historical wind patterns, we store and replay only a small portion of historical data named episodic memory. A simple but efficient strategy is designed to constitute episodic memory and thus relieve the computation burden. Experiments conducted on two real-world datasets demonstrate the superiority of our method over existing approaches. Particularly, WindTrans surpasses state-of-the-art methods by up to 36.7%, 29.3% and 13.3% improvement in MAPE measure for 1 hour ahead prediction on 10-minute, 5-minute, and 1-minute-based tasks, respectively. Furthermore, via our continual learning scheme, the model retains competitive performance with only 6.9% datum stored and retrained on.
Chen Liu 0034, Shibo He, Haoyu Liu 0002, Jiming Chen 0001, Hairong Dong 0001
IEEE Trans. Intell. Transp. Syst.3
2024 Label-Free Multivariate Time Series Anomaly Detection
abstract
Anomaly detection in multivariate time series has been widely studied in one-class classification (OCC) setting. The training samples in this setting are assumed to be normal. In more practical situations, it is difficult to guarantee that all samples are normal. Meanwhile, preparing a completely clean training dataset is costly and laborious. Such a case may degrade the performance of OCC-based anomaly detection methods which fit the training distribution as the normal distribution. To overcome this limitation, in this paper, we propose MTGFlow, an unsupervised anomaly detection approach for Multivariate Time series anomaly detection via dynamic Graph and entity-aware normalizing Flow. MTGFlow first estimates the density of the entire training samples and then identifies anomalous instances based on the density of the test samples within the fitted distribution. This relies on a widely accepted assumption that anomalous instances exhibit more sparse densities than normal ones, with no reliance on the clean training dataset. However, it is intractable to directly estimate the density due to the complex dependencies among entities and their diverse inherent characteristics, not to mention detecting anomalies based on the estimated distribution. In order to address these problems, we utilize the graph structure learning model to learn interdependent and evolving relations among entities, which effectively captures the complex and accurate distribution patterns of multivariate time series. In addition, our approach incorporates the unique characteristics of individual entities by employing an entity-aware normalizing flow. This enables us to represent each entity as a parameterized normal distribution. Furthermore, considering that some entities present similar characteristics, we propose a cluster strategy that capitalizes on the commonalities of entities with similar characteristics, resulting in more precise and detailed density estimation. We refer to this cluster-aware extension as MTGFlow_cluster. Extensive experiments are conducted on six widely used benchmark datasets, in which MTGFlow and MTGFlow_cluster demonstrate their superior detection performance.
Qihang Zhou, Shibo He, Haoyu Liu 0002, Jiming Chen 0001, Wenchao Meng
IEEE Trans. Knowl. Data Eng.3
2023 Detecting Multivariate Time Series Anomalies with Zero Known Label
abstract
Multivariate time series anomaly detection has been extensively studied under the one-class classification setting, where a training dataset with all normal instances is required. However, preparing such a dataset is very laborious since each single data instance should be fully guaranteed to be normal. It is, therefore, desired to explore multivariate time series anomaly detection methods based on the dataset without any label knowledge. In this paper, we propose MTGFlow, an unsupervised anomaly detection approach forMultivariate Time series anomaly detection via dynamic Graph and entityaware normalizing Flow, leaning only on a widely accepted hypothesis that abnormal instances exhibit sparse densities than the normal. However, the complex interdependencies among entities and the diverse inherent characteristics of each entity pose significant challenges to density estimation, let alone to detect anomalies based on the estimated possibility distribution. To tackle these problems, we propose to learn the mutual and dynamic relations among entities via a graph structure learning model, which helps to model the accurate distribution of multivariate time series. Moreover, taking account of distinct characteristics of the individual entities, an entity-aware normalizing flow is developed to describe each entity into a parameterized normal distribution, thereby producing fine-grained density estimation. Incorporating these two strategies, MTGFlow achieves superior anomaly detection performance. Experiments on five public datasets with seven baselines are conducted, MTGFlow outperforms the SOTA methods by up to 5.0 AUROC%.
Qihang Zhou, Jiming Chen 0001, Haoyu Liu 0002, Shibo He, Wenchao Meng
AAAI3
2023 InstanT: Semi-supervised Learning with Instance-dependent Thresholds
abstract
Semi-supervised learning (SSL) has been a fundamental challenge in machine learning for decades. The primary family of SSL algorithms, known as pseudo-labeling, involves assigning pseudo-labels to confident unlabeled instances and incorporating them into the training set. Therefore, the selection criteria of confident instances are crucial to the success of SSL. Recently, there has been growing interest in the development of SSL methods that use dynamic or adaptive thresholds. Yet, these methods typically apply the same threshold to all samples, or use class-dependent thresholds for instances belonging to a certain class, while neglecting instance-level information. In this paper, we propose the study of instance-dependent thresholds, which has the highest degree of freedom compared with existing methods. Specifically, we devise a novel instance-dependent threshold function for all unlabeled instances by utilizing their instance-level ambiguity and the instance-dependent error rates of pseudo-labels, so instances that are more likely to have incorrect pseudo-labels will have higher thresholds. Furthermore, we demonstrate that our instance-dependent threshold function provides a bounded probabilistic guarantee for the correctness of the pseudo-labels it assigns.
Runze Wu 0001, Haoyu Liu 0002, Jun Yu 0001, Xun Yang 0001, Bo Han 0003, Tongliang Liu
NeurIPS3
2023 Pull & Push: Leveraging Differential Knowledge Distillation for Efficient Unsupervised Anomaly Detection and Localization
abstract
Recently, much attention has been paid to segmenting subtle unknown defect regions by knowledge distillation in an unsupervised setting. Most previous studies concentrated on guiding the student network to learn the same representations on the normality, neglecting the different behaviors of the abnormality. This leads to a high probability of false detection of subtle defects. To address such an issue, we propose to push representations on abnormal areas of the teacher and student network as far as possible while pulling representations on normal areas as close as possible. Based on this idea, we design an efficient teacher-student model for anomaly detection and localization, which maximizes pixel-wise discrepancies for anomalous regions approximated by data augmentation and simultaneously minimizes discrepancies for pixel-wise normal regions between these two networks. The explicit differential knowledge distillation enlarges the margin between normal representations and abnormal ones in favour of discriminating them. Then, the appropriate small student network is not only efficient, but more importantly, helps inhibit the generalization ability of anomalous patterns when learning normal patterns, facilitating the precise decision boundary. The experimental results on the MVTec AD, Fashion-MNIST, and CIFAR-10 datasets demonstrate that our proposed method achieves better performance than current state-of-the-art (SOTA) approaches. Especially, For the MVTec AD dataset with high resolution images, we achieve 98.1 AUROC% and 93.6 AUPRO% in anomaly localization, outperforming knowledge distillation based SOTA methods by 1.1 AUROC% and 1.5 AUPRO% with a lightweight model.
Qihang Zhou, Shibo He, Haoyu Liu 0002, Tao Chen 0003, Jiming Chen 0001
IEEE Trans. Circuits Syst. Video Technol.3
2022 Generalized Global Ranking-Aware Neural Architecture Ranker for Efficient Image Classifier Search
abstract
Neural Architecture Search (NAS) is a powerful tool for automating effective image processing DNN designing. The ranking has been advocated to design an efficient performance predictor for NAS. The previous contrastive method solves the ranking problem by comparing pairs of architectures and predicting their relative performance. However, it only focuses on the rankings between two involved architectures and neglects the overall quality distributions of the search space, which may suffer generalization issues. A predictor, namely Neural Architecture Ranker (NAR) which concentrates on the global quality tier of specific architecture, is proposed to tackle such problems caused by the local perspective. The NAR explores the quality tiers of the search space globally and classifies each individual to the tier they belong to according to its global ranking. Thus, the predictor gains the knowledge of the performance distributions of the search space which helps to generalize its ranking ability to the datasets more easily. Meanwhile, the global quality distribution facilitates the search phase by directly sampling candidates according to the statistics of quality tiers, which is free of training a search algorithm, e.g., Reinforcement Learning (RL) or Evolutionary Algorithm (EA), thus it simplifies the NAS pipeline and saves the computational overheads. The proposed NAR achieves better performance than the state-of-the-art methods on two widely used datasets for NAS research. On the vast search space of NAS-Bench-101, the NAR easily finds the architecture with top 0.01 performance only by sampling. It also generalizes well to different image datasets of NAS-Bench-201, i.e., CIFAR-10, CIFAR-100, and ImageNet-16-120 by identifying the optimal architectures for each of them.
Bicheng Guo, Tao Chen 0003, Shibo He, Haoyu Liu 0002, Lilin Xu, Peng Ye 0006, Jiming Chen 0001
ACM Multimedia4
2022 Toward Optimal Deployment for Full-View Point Coverage in Camera Sensor Networks
abstract
Recent years have witnessed the fast proliferation of camera sensors networks (CSNs) in numerous Internet of Things (IoT) applications. In order to a capture distinct image of targets from interesting directions, we leverage a special type of coverage called full-view coverage. Full-view coverage guarantees to obtain the images of a point from every direction, whereas it demands much more sensors than a conventional coverage. To this end, we investigate the problem of deploying the minimum number of rotatable camera sensors to achieve the full-view coverage of a set of target points, namely, optimal deployment for the full-view point coverage (OFP) problem. In this work, camera sensors are capable of rotating freely with infinite orientations, thus not only the deployment locations but also the orientations for each camera sensor are required to be optimized. To tackle this challenging problem, we first prove that the OFP problem is NP-hard. Then, we propose two approximation algorithms—iterative screening algorithm (ISA) and improved ISA (IISA) to solve the OFP. We further perform extensive simulations and conduct physical testings to demonstrate the superiority and effectiveness of our proposed solutions. Experimental results show that IISA can generally reduce the total number of required camera sensors by more than 20% compared with the state-of-the-art work.
Kun Shi 0003, Shuxian Liu, Chao Li 0062, Haoyu Liu 0002, Shibo He, Qi Zhang 0066, Jiming Chen 0001
IEEE Internet Things J.4
2022 Towards Automatic Root Cause Diagnosis of Persistent Packet Loss in Cloud Overlay Network
abstract
Persistent packet loss in the cloud-scale overlay network severely compromises tenant experiences. Cloud providers are keen to diagnose such problems efficiently. However, existing work is either designed for the physical network or insufficient to present the concrete reason of packet loss. We propose to record and analyze the on-site forwarding condition of packets during packet-level tracing. The cloud-scale overlay network presents great challenges to achieve this goal with its high network complexity, multi-tenant nature, and diversity of root causes. To address these challenges, we present VTrace, an automatic diagnostic system for persistent packet loss over the cloud-scale overlay network. Utilizing the “fast path-slow path” structure of virtual forwarding devices (VFDs), e.g., vSwitches, VTrace installs several “coloring-matching-logging” rules in VFDs to selectively track the target packets and inspect them in depth. The detailed forwarding situation at each hop is logged and then assembled to perform analysis with an efficient path reconstruction scheme. Experiments are conducted to demonstrate VTrace’s low overhead and quick response. Besides, based on the idea “coloring-matching-counting”, VTrace can be easily extended toVTrace-statsto identify the culprit device for transient packet loss. We share experiences of how VTrace andVTrace-statsefficiently work after deploying them in Alibaba Cloud for years.
Chongrong Fang, Haoyu Liu 0002, Mao Miao, Lei Wang 0005, Wansheng Zhang, Daxiang Kang, Biao Lyu, Shunmin Zhu, Peng Cheng 0001, Jiming Chen 0001
IEEE/ACM Trans. Netw.2
2021 A survey of cloud network fault diagnostic systems and tools
abstract
Recently, cloud computing has become a vital part that supports people’s normal lives and production. However, accompanied by the increasing complexity of the cloud network, failures constantly keep coming up and cause huge economic losses. Thus, to guarantee the cloud network performance and prevent execrable effects caused by failures, cloud network diagnostics has become of great interest for cloud service providers. Due to the characteristics of cloud network (e.g., virtualization and multi-tenancy), transplanting traditional network diagnostic tools to the cloud network face several difficulties. Additionally, many existing tools cannot solve problems in the cloud network. In this paper, we summarize and classify the state-of-the-art technologies of cloud diagnostics which can be used in the production cloud network according to their features. Moreover, we analyze the differences between cloud network diagnostics and traditional network diagnostics based on the characteristics of the cloud network. Considering the operation requirements of the cloud network, we propose the points that should be cared about when designing a cloud network diagnostic tool. Also, we discuss the challenges that cloud network diagnostics will face in future development.
Yining Qi, Chongrong Fang, Haoyu Liu 0002, Daxiang Kang, Biao Lyu, Peng Cheng 0001, Jiming Chen 0001
Frontiers Inf. Technol. Electron. Eng.3
2021 Short-Term Strong Wind Risk Prediction for High-Speed Railway
abstract
Running at a fast speed, the high-speed train is prone to be interrupted by the surrounding strong wind. To ensure the safety of the trains, an effective approach is to deploy anemometers alongside the railway, such that the real-time and short-term predicted wind speed can be reported, and be further used by dispatchers to take protective actions in advance. However, in certain situations, the solely predicted wind speed is not informative enough to describe the wind status. It is difficult to tell if a strong wind incident could happen when the predicted wind speed is slightly lower than the strong wind threshold. We take the first attempt to predict the strong wind risk alongside the high-speed railway (HSR). A new model, called Multiple Attention Layer based Multi-Instance Learning (MAL-MIL), is proposed to address this problem. The key idea is to estimate the possibility that the actual wind speed exceeds the threshold conditionally on the predicted wind status. Based on attention mechanisms and long-short term memory network, the model can firstly generate deep representations of the future wind status. Then, though there is a lack of the risk ground truth, the multi-instance learning process facilitates the training procedure so that the relationships between these deep representations and the strong wind incidents could be quantified. Furthermore, considering the practicality of the model, we also design a result justification module to explain the reported risk. The superior performance is finally verified based on a real-world dataset.
Haoyu Liu 0002, Chen Liu 0034, Shibo He, Jiming Chen 0001
IEEE Trans. Intell. Transp. Syst.1
2020 LP-Explain: Local Pictorial Explanation for Outliers
abstract
Outlier detection is of vital importance for various fields and applications. Existing works mainly focus on identifying outliers from underlying datasets, while how to provide sense-making explanations is largely ignored. In this paper, we propose to visualize data points in a set of scatter plots on two-dimensional (2-D) feature spaces that can provide meaningful explanations about the outlying behavior of outliers. Data are typically multidimensional and the number of 2-D combinations could be huge. Also, outliers may have diverse characteristics, and thus the global scatter plots containing all of outliers may degrade the explanation effectiveness for those outliers having idiosyncratic abnormal 2-D spaces. To address this problem, we propose a new outlier explanation approach, called LP-Explain, which tries to identify the set of best Local Pictorial explanations (defined as the scatter plots in the 2-D space of the feature pairs) that can Explain the behavior for cluster of outliers. We first define an effective measure to quantify the similarity between outliers, and then cluster outliers into different groups based on their abnormal feature pairs. We then propose to weigh the importance of feature pairs within each cluster through a multi-task learning framework to select the set of top feature pairs that best explain various outlier clusters. By adjusting a user-defined parameter indicating the “localization level”, the proposed method can attain both global and local results for the explanation of the outliers. 2-D visual explanations can be plotted for the top-weighted feature pairs of each cluster. We conduct experiments on various public datasets, which show that the proposed approach can provide more meaningful explanations about the outlying behavior in a dataset.
Haoyu Liu 0002, Fenglong Ma, Yaqing Wang 0001, Shibo He, Jiming Chen 0001, Jing Gao 0004
ICDM1
2020 RAIN: Towards Real-Time Core Devices Anomaly Detection Through Session Data in Cloud Network
abstract
Core devices form the critical components of the cloud network and provide service to multiple tenants simultaneously. The anomalies that happened in core devices impact network availability of a large number of users, meanwhile, lead to the degradation of cloud providers’ profits. However, direct monitoring of core devices needs to deploy massive heartbeat checking tools on numerous related components, which will be extremely laborious. In this paper, we deploy RAIN to reduce the number of devices that need to be detailed investigated for anomalies. The session traffic data among core devices and served virtual machines are utilized to conduct the analyzing. To guarantee near real-time monitoring, RAIN is designed as a two-step structure and incorporating four feature-based detection methods. RAIN has been deployed in Alibaba’s production cloud network for over 6 months and is analyzing terabytes of traffic flow metrics per day.
Haoyu Liu 0002, Chongrong Fang, Yining Qi, Shaozhe Wang, Daxiang Kang, Biao Lyu, Peng Cheng 0001, Jiming Chen 0001
NOMS1
2020 VTrace: Automatic Diagnostic System for Persistent Packet Loss in Cloud-Scale Overlay Network
abstract
Persistent packet loss in the cloud-scale overlay network severely compromises tenant experiences. Cloud providers are keen to automatically and quickly determine the root cause of such problems. However, existing work is either designed for the physical network or insufficient to present the concrete reason of packet loss. In this paper, we propose to record and analyze the on-site forwarding condition of packets during packet-level tracing. The cloud-scale overlay network presents great challenges to achieve this goal with its high network complexity, multi-tenant nature, and diversity of root causes. To address these challenges, we present VTrace, an automatic diagnostic system for persistent packet loss over the cloud-scale overlay network. Utilizing the "fast path-slow path" structure of virtual forwarding devices (VFDs), e.g., vSwitches, VTrace installs several "coloring, matching and logging" rules in VFDs to selectively track the packets of interest and inspect them in depth. The detailed forwarding situation at each hop is logged and then assembled to perform analysis with an efficient path reconstruction scheme. Experiments are conducted to demonstrate VTrace's low overhead and quick responsiveness. We share experiences of how VTrace efficiently resolves persistent packet loss issues after deploying it in Alibaba Cloud for over 20 months.
Chongrong Fang, Haoyu Liu 0002, Mao Miao, Lei Wang 0005, Wansheng Zhang, Daxiang Kang, Biao Lyu, Peng Cheng 0001, Jiming Chen 0001
SIGCOMM2
2019 Orientation Optimization for Full-View Coverage Using Rotatable Camera Sensors
abstract
Recently, full-view coverage has been introduced to capture intruders from multiple directions in the camera sensor networks. It is more efficient than traditional coverage in identifying the intruders. However, full-view coverage typically calls for a large number of camera sensors. Hence, we exploit limited mobility or orientation to improve the performance of full-view coverage since camera sensors typically can rotate to cover more areas without being relocated after installation. Observing that target points may not be full-view covered constantly due to the sensor rotation, we emphasize the importance of the fairness-based coverage maximization problem, i.e., how to schedule the orientations of camera sensors to maximize the minimum cumulative full-view coverage time of target points. To solve this issue, we first try to reduce the dimension space of orientations by dividing the orientation space into a set of discrete directions. We then study how to select the minimum number of sensing regions that camera sensors should rotate to cover in order to ensure the full-view coverage of all target points. Next, we unveil the relationship between the full-view coverage and target points, which are spatially correlated. Based on these results, we devise a centralized algorithm to solve the problem based on “largest demand first serve” principle, by which the target points with less cumulative full-view coverage time will be preferentially selected to be full-view covered with a higher probability. We further design a distributed solution as a counterpart of the centralized algorithm. Extensive simulations are presented to show the performances of the proposed algorithms. Results show that exploiting limited mobility of sensor rotation has good potential in promoting the efficiency and reducing the cost of ensuring full-view coverage.
Jiming Chen 0001, Haoyu Liu 0002, Qi Zhang 0066, Shibo He
IEEE Internet Things J.2