VLDB 2026 Research / reviewers in the wild / expert
Yongqiang Yang
dblp:23/10161
· DBLP profile ↗
30ranked-venue papers
1as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 9 · 9 since 2021Systems, architecture and hardware · 8 · 7 since 2021Computer networks · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Security and privacy · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Misclassification-Aware Robust Learning from Multiple Human Labelers (Student Abstract)abstractAdversarial training is an effective technique for enhancing the robustness of deep neural networks (DNNs). Prior research shows that misclassified examples influence final adversarial robustness much more than correctly classified examples. Ignoring this difference during training can hurt model performance. In crowdsourcing, varying annotator expertise causes noisy, inconsistent labels. As a result, it is hard to distinguish misclassified and correctly classified examples using only provided annotations. Thus, how to use the reliability and discrepancy between these example types to improve robustness within adversarial learning remains a critical but underexplored issue. In this work, we first explore how misclassified and correctly classified examples affect learning from crowds (LFC) in adversarial environments. Then, we formulate the problem of misclassification-aware robust learning from multiple human labelers as a bilevel min-max problem. After that, we introduce MALC, a new approach to make classifiers more robust to adversarial examples via iterative adversarial example generation and parameter estimation. We conduct an extensive evaluation of the proposed MALC, showing that MALC can outperform the state-of-the-art LFC methods in both white-box and black-box settings. Zuoyuehe Wang, Chicheng Ma, Lei Chai, Yongqiang Yang, Jingzheng Li |
AAAI | 5 |
| 2026 | LLMGuard: Multi-Agent Fault Diagnosis for Reliable Language-Model-as-a-Service
Yuedong Zhong, Guangba Yu, QunChao Fu, Yongqiang Yang, Michael R. Lyu |
DSN | 7 |
| 2026 | Rearchitecting Programmable Networks For In-Network Computing: From Hardware To LanguageabstractIn-network computing (INC) offers orders-of-magnitude performance gains for various applications. However, the existing pipeline-based switch architecture, though well-suited for stateless packet forwarding, imposes three constraints on stateful INC applications: (i) control-plane table management, (ii) scarce memory resources under per-stage layout, and (iii) a strict memory access model. To address these constraints, this paper provides a full-stack solution, covering a programmable chipset Solar-NP, a programming language NPC, and a complete toolchain XuanWu. First, Solar-NP adopts a run-to-completion-based (RTC-based) architecture with three new hardware features: data-plane table management, a hierarchical memory pool, and built-in data structures with atomic access guarantees. Second, to fully exploit the hardware features, our language NPC introduces a new abstraction, namely operation-action table (OAT), that allows table manipulations in the data plane. Finally, our XuanWu toolchain provides various utility software covering the complete workflow of developing, testing, debugging, and deployment. To show the power of our solution, we implement three types of INC cases, including in-network control, in-network telemetry, and in-network storage. Experimental results demonstrate that our solution allows more operations to be offloaded and achieves higher performance than today's pipeline-based solutions. Haifeng Sun 0002, Taixu Tian, Jinbo Sun, Jintao He, Qun Huang 0001, Luyou He, Xiangcan Xu, Junyi Guo, Yongqiang Yang |
EuroSys | 14 |
| 2026 | Suika: Efficient and High-quality Re-scheduling of 3D-parallelized LLM Training Jobs in Shared ClustersabstractLarge Language Models (LLMs) are usually trained with 3D (data, tensor, and pipeline) parallelism—in shared GPU clusters where the available resources are highly dynamic. Rescheduling the idle resources to ongoing jobs can help improve cluster utilization, but doing so for 3D-parallelized training jobs suffers large overheads in performance modeling, decision making, and redeployment. We present Suika, a cluster training system that supports efficient and high-quality resource rescheduling for 3D-parallelized LLM training jobs. Suika holistically addresses the complexity challenges by exploiting the incremental nature of rescheduling. For performance modeling, it builds an accurate performance estimator with non-disruptive online profiling. For decision-making, it employs topology-aware sorting and an expand-and-balance algorithm to reduce the complexity of resource allocation and job parallelization, without compromising decision quality. Suika further integrates a device-to-device redeployment method to leverage the overlapping nature of incremental reconfiguration for overhead reduction. Experiments on 64-GPU physical cluster and 1024-GPU simulated cluster show that, Suika achieves 1.29 ~ 1.31× reduction in average JCT compared to state-of-the-art schedulers. Chen Chen 0067, Chunyu Xue, Qizhen Weng 0001, Zeren Li, Xuqi Zhu, Yongqiang Yang, Quan Chen 0002, Minyi Guo |
EuroSys | 9 |
| 2026 | Meteor: High-Performance Control Message Delivery for Large-Scale CloudsabstractVirtual private clouds (VPCs) play a critical role in providing secure and isolated network environments for web services. However, with the growing number and size of VPCs, efficiently delivering control messages from the control plane to the data plane has become a major concern for cloud vendors. Existing end-to-end transmission solutions (e.g., RPC) will result in substantial overhead in the control plane, while message-oriented middleware-based solutions (e.g., message queue) will lead to high data plane overhead. To address this issue, we design Meteor, a high-performance control message delivery system for large-scale clouds. Specifically, Meteor combines an RPC path with a message queue (MQ) path and employs an auto dual-path switching mechanism to minimize the message delivery latency. Additionally, we propose a VPC-based message delivery and filtering scheme for the MQ path to reduce data plane overhead. We also design a delivery robustness guarantee mechanism to ensure the reachability and consistency of control messages. Meteor has been thoroughly tested with up to 100k container instances. Evaluation results show that Meteor decreases the message delivery latency by 48.8% and reduces the overhead by about 50% in real-world scenarios, compared with state-of-the-art solutions. Gongming Zhao, Baoqing Wang, Min Chen 0033, Hongli Xu 0001, Jiawei Liu 0007, Xuwei Yang, Liguang Xie, Yongqiang Yang |
WWW | 8 |
| 2026 | KPIRoot+: An efficient integrated framework for anomaly detection and root cause analysis in large-scale cloud systems
Wenwei Gu, Renyi Zhong, Guangba Yu, Xinying Sun, Jinyang Liu 0002, Yintong Huo, Zhuangbin Chen, Jianping Zhang 0002, Jiazhen Gu, Yongqiang Yang, Michael R. Lyu |
Empir. Softw. Eng. | 10 |
| 2026 | New fusion methods using a weighted L1 conflict measure for reliability assessment with multi-credibility evidence
Yongqiang Yang, Bin Suo, Ruxin Yin, Yanrui He |
Inf. Sci. | 1 |
| 2026 | Network Specification Mining With High Fidelity, Scalability, and ReadabilityabstractNetwork specification, which describes what an existing network is designed for, can help operators better understand and manage their networks, and is a critical pre-condition for network verification and synthesis tools to work. Existing tools for specification mining either cannot scale to large networks, or scale by sacrificing fidelity. Moreover, the specification contains a huge number of low-level intents (e.g., tens of thousands of pairwise reachability), making it hard for operators to read. To this end, this paper presentsNetMiner, which can mine specification from network configurations, with high scalability, fidelity, and easier to read. The key idea ofNetMineris to faithfully simulate the network routing and forwarding behaviors with control plane simulators and data plane verifiers, so as to achieve high fidelity. Meanwhile,NetMinerimproves the scalability by identifying relevant failure scenarios, and aggregating them to significantly reduce the number of needed simulations. Moreover,NetMinerclusters similar low-level intents into a high-level intent, to make the specification more concise and easier to read. Experiments using real configurations from a large cloud service provider and synthetic configurations show thatNetMinercan mine specification$10\times $faster, and reduce the number of intents by$100\times $, compared to state-of-the-art tools. Ning Kang 0003, Peng Zhang 0011, Hao Li 0011, Sisi Wen, Chaoyang Ji, Yongqiang Yang |
IEEE Trans. Netw. | 6 |
| 2025 | ADAMAS: Adaptive Domain-Aware Performance Anomaly Detection in Cloud Service SystemsabstractA common practice in the reliability engineering of cloud services involves the collection of monitoring metrics, followed by comprehensive analysis to identify performance issues. However, existing methods often fall short of detecting diverse and evolving anomalies across different services. More-over, there exists a significant gap between the technical and business interpretation of anomalies, i.e., a detected anomaly may not have an actual impact on system performance or user experience. To address these challenges, we propose ADAMAS, an adaptive AutoML-based anomaly detection framework aiming to achieve practical anomaly detection in production cloud systems. To improve the ability to detect cross-service anomalies, we design a novel unsupervised evaluation function to facilitate the automatic searching of the optimal model structure and parameters. ADAMAS also contains a lightweight human-in-the-loop design, which can efficiently incorporate expert knowledge to adapt to the evolving anomaly patterns and bridge the gap between predicted anomalies and actual business exceptions. Fur-thermore, through monitoring the rate of mispredicted anomalies, ADAMAS proactively re-configures the optimal model, forming a continuous loop of system improvement. Extensive evaluation on one public and two industrial datasets shows that ADAMAS outperforms all baseline models with a 0.891 F1-score. The ablation study also proves the effectiveness of the evaluation function design and the incorporation of expert knowledge. Wenwei Gu, Jiazhen Gu, Jinyang Liu 0002, Zhuangbin Chen, Jianping Zhang 0002, Jinxi Kuang, Yongqiang Yang, Michael R. Lyu |
ICSE | 8 |
| 2025 | Identifying Performance Issues in Cloud Service Systems Based on Relational-Temporal FeaturesabstractCloud systems, typically comprised of various components (e.g., microservices), are susceptible to performance issues, which may cause service-level agreement violations and financial losses. Identifying performance issues is thus of paramount importance for cloud vendors. In current practice, crucial metrics, i.e., Key Performance Indicators (KPIs), are monitored periodically to provide insight into the operational status of components. Identifying performance issues is often formulated as an anomaly detection problem, which is tackled by analyzing each metric independently. However, this approach overlooks the complex dependencies existing among cloud components. Some graph neural network-based methods take both temporal and relational information into account; however, the correlation violations in the metrics that serve as indicators of underlying performance issues are difficult for them to identify. Furthermore, a large volume of components in a cloud system results in a vast array of noisy metrics. This complexity renders it impractical for engineers to fully comprehend the correlations, making it challenging to identify performance issues accurately. To address these limitations, we propose Identifying Performance Issues based on Relational-Temporal Features (ISOLATE), a learning-based approach that leverages both the relational and temporal features of metrics to identify performance issues. In particular, it adopts a graph neural network with attention to characterizing the relations among metrics and extracts long-term and multi-scale temporal patterns using a GRU and a convolution network, respectively. The learned graph attention weights can be further used to localize the correlation-violated metrics. Moreover, to relieve the impact of noisy data, ISOLATE utilizes a Positive Unlabeled (PU) Learning strategy that tags pseudo-labels based on a small portion of confirmed negative examples. Extensive evaluation on both public and industrial datasets shows that ISOLATE outperforms all baseline models with 0.945 F1 score and 0.920 Hit rate@3. The ablation study also proves the effectiveness of the relational-temporal features and the PU-Learning strategy. Furthermore, we share the success stories of leveraging ISOLATE to identify performance issues in Huawei Cloud, which demonstrates its superiority in practice. Wenwei Gu, Jinyang Liu 0002, Zhuangbin Chen, Jianping Zhang 0002, Yuxin Su 0001, Jiazhen Gu, Zengyin Yang, Yongqiang Yang, Michael R. Lyu |
ACM Trans. Softw. Eng. Methodol. | 9 |
| 2024 | FEDGE: An Interference-Aware QoS Prediction Framework for Black-Box Scenario in IaaS Clouds with Domain GeneralizationabstractPublic cloud providers embrace multi-tenancy as a strategy to enhance the utilization and efficiency of resources. However, co-located virtual machines (VMs) suffer from qualityof-service (QoS) degradation caused by shared resource interference. Existing solutions for predicting QoS degradation often rely on the assumption of online access to application-level information. However, in a production environment, this assumption proves invalid as the VMs are black boxes to the providers. This intrinsic characteristic of the IaaS cloud necessitates the prediction model to generalize to unfamiliar applications and imposes specific criteria on the monitorable metrics.To meet the black-box scenario under Infrastructure as a Service (IaaS) cloud computing, we present a novel framework, FEDGE, that can predict interference-aware QoS (IA-QoS) of co-located VMs using only low-level monitorable metrics before migration. Specifically, FEDGE utilizes a stochastic gates layer to select the most informative features from the high-dimensional resource and hardware metrics, which helps to reduce the monitoring overhead. Furthermore, we design a multi-domain MMD-based adversarial denoising autoencoder to regularize the learned hidden representations and prevent over-fitting on the source domains. Next, we employ a multi-layer perceptron (MLP) to accurately predict complex QoS degradation using the learned representations with domain generalization. Experimental results demonstrate that FEDGE outperforms other state-of-the-art methods in terms of both generalizability and effectiveness. Yunlong Cheng, Xiuqi Huang, Zifeng Liu, Jiadong Chen, Xiaofeng Gao 0001, Yongqiang Yang |
IPDPS | 7 |
| 2024 | KPIRoot: Efficient Monitoring Metric-based Root Cause Localization in Large-scale Cloud SystemsabstractTo ensure the reliability of cloud systems, their run-time status reflecting the service quality is periodically monitored with monitoring metrics, i.e., KPIs (key performance indicators). When performance issues happen, root cause localization pinpoints the specific KPIs that are responsible for the degradation of overall service quality, facilitating prompt problem diagnosis and resolution. To this end, existing methods generally locate root-cause KPIs by identifying the KPIs that exhibit a similar anomalous trend to the overall service performance. While straightforward, solely relying on the similarity calculation may be ineffective when dealing with cloud systems with complicated interdependent services. Recent deep learning-based methods offer improved performance by modeling these intricate dependencies. However, their high computational demand often hinders their ability to meet the efficiency requirements of industrial applications. Furthermore, their lack of interpretability further restricts their practicality. To overcome these limitations, we propose KPIRoot, an effective and efficient method for root cause localization integrating both advantages of similarity analysis and causality analysis, where similarity measures the trend alignment of KPI and causality measures the sequential order of variation of KPI. Furthermore, we leverage symbolic aggregate approximation to produce a more compact representation for each KPI, enhancing the overall analysis efficiency of the approach. The experimental results show that KPIRoot outperforms seven state-of-the-art baselines by 7.9%~28.3%, while time cost is reduced by 56.9%. Moreover, we share our experience of deploying KPIRoot in the production environment of a large-scale cloud provider Cloud ${\mathcal{H}^{\ast}}$. Wenwei Gu, Xinying Sun, Jinyang Liu 0002, Yintong Huo, Zhuangbin Chen, Jianping Zhang 0002, Jiazhen Gu, Yongqiang Yang, Michael R. Lyu |
ISSRE | 8 |
| 2024 | MicroRes: Versatile Resilience Profiling in Microservices via Degradation Dissemination IndexingabstractMicroservice resilience, the ability of microservices to recover from failures and continue providing reliable and responsive services, is crucial for cloud vendors. However, the current practice relies on manually configured rules specific to a certain microservice system, resulting in labor-intensity and flexibility issues, given the large scale and high dynamics of microservices. A more labor-efficient and versatile solution is desired. Our insight is that resilient deployment can effectively prevent the dissemination of degradation from system performance metrics to user-aware metrics, and the latter directly affects service quality. In other words, failures in a non-resilient deployment can impact both types of metrics, leading to user dissatisfaction. With this in mind, we propose MicroRes, the first versatile resilience profiling framework for microservices via degradation dissemination indexing. MicroRes first injects failures into microservices and collects available monitoring metrics. Then, it ranks the metrics according to their contributions to the overall service degradation. It produces a resilience index by how much the degradation is disseminated from system performance metrics to user-aware metrics. Higher degradation dissemination indicates lower resilience. We evaluate MicroRes on two open-source and one industrial microservice system. The experiments show MicroRes' efficient and effective resilience profiling of microservices. We also showcase MicroRes' practical usage in production. Cheryl Lee, Jiacheng Shen, Yuxin Su 0001, Yongqiang Yang, Michael R. Lyu |
ISSTA | 6 |
| 2023 | Network Specification Mining with High Fidelity and ScalabilityabstractNetwork specification, which describes what an existing network is designed for, can help operators better understand and manage their networks, and is a critical pre-condition for network verification and synthesis tools to work. Mining specification with existing tools either cannot scale to large networks, or scale better at a cost of sacrificing fidelity. This paper presents NetMiner, which can simultaneously achieve high scalability and fidelity. The key idea of NetMiner is to use off-the-shelf network simulators to compute routes, and then check properties with data plane verifiers, so as to achieve a high fidelity. At the same time, NetMiner improves the scalability by identifying relevant failure scenarios, and aggregating them to significantly reduce the number of needed simulations. This process is solely based on the routes returned by the simulators and therefore preserving fidelity. Experiments using real configurations from a large cloud service provider and synthetic configurations show that NetMiner is about 10 times faster than the state-of-the-art. Ning Kang 0001, Sisi Wen, Chaoyang Ji, Yongqiang Yang |
ICNP | 6 |
| 2023 | Heterogeneous Anomaly Detection for Software Systems via Semi-supervised Cross-modal AttentionabstractPrompt and accurate detection of system anomalies is essential to ensure the reliability of software systems. Unlike manual efforts that exploit all available run-time information, existing approaches usually leverage only a single type of monitoring data (often logs or metrics) or fail to make effective use of the joint information among different types of data. Consequently, many false predictions occur. To better understand the manifestations of system anomalies, we conduct a systematical study on a large amount of heterogeneous data, i.e., logs and metrics. Our study demonstrates that logs and metrics can manifest system anomalies collaboratively and complementarily, and neither of them only is sufficient. Thus, integrating heterogeneous data can help recover the complete picture of a system's health status. In this context, we propose Hades, the first end-to-end semi-supervised approach to effectively identify system anomalies based on heterogeneous data. Our approach employs a hierarchical architecture to learn a global representation of the system status by fusing log semantics and metric patterns. It captures discriminative features and meaningful interactions from heterogeneous data via a cross-modal attention module, trained in a semi-supervised manner. We evaluate Hades extensively on large-scale simulated data and datasets from Huawei Cloud. The experimental results present the effectiveness of our model in detecting system anomalies. We also release the code and the annotated dataset for replication and future research. Cheryl Lee, Zhuangbin Chen, Yuxin Su 0001, Yongqiang Yang, Michael R. Lyu |
ICSE | 5 |
| 2023 | Black-Box Data Poisoning Attacks on CrowdsourcingabstractUnderstanding the vulnerability of label aggregation against data poisoning attacks is key to ensuring data quality in crowdsourced label collection. State-of-the-art attack mechanisms generally assume full knowledge of the aggregation models while failing to consider the flexibility of malicious workers in selecting which instances to label. Such a setup limits the applicability of the attack mechanisms and impedes further improvement of their success rate. This paper introduces a black-box data poisoning attack framework that finds the optimal strategies for instance selection and labeling to attack unknown label aggregation models in crowdsourcing. We formulate the attack problem on top of a generic formalization of label aggregation models and then introduce a substitution approach that attacks a substitute aggregation model in replacement of the unknown model. Through extensive validation on multiple real-world datasets, we demonstrate the effectiveness of both instance selection and model substitution in improving the success rate of attacks. Yongqiang Yang, Dingqi Yang, Hailong Sun 0001 |
IJCAI | 2 |
| 2023 | Alioth: A Machine Learning Based Interference-Aware Performance Monitor for Multi-Tenancy Applications in Public CloudabstractMulti-tenancy in public clouds may lead to co-location interference on shared resources, which possibly results in performance degradation of cloud applications. Cloud providers want to know when such events happen and how serious the degradation is, to perform interference-aware migrations and alleviate the problem. However, virtual machines (VM) in Infrastructure-as-a-Service public clouds are black boxes to providers, where application-level performance information cannot be acquired. This makes performance monitoring intensely challenging as cloud providers can only rely on low-level metrics such as CPU usage and hardware counters.We propose a novel machine learning framework, Alioth, to monitor the performance degradation of cloud applications. To feed the data-hungry models, we first elaborate interference generators and conduct comprehensive co-location experiments on a testbed to build Alioth-dataset which reflects the complexity and dynamicity in real-world scenarios. Then we construct Alioth by (1) augmenting features via recovering low-level metrics under no interference using denoising auto-encoders, (2) devising a transfer learning model based on domain adaptation neural network to make models generalize on test cases unseen in offline training, and (3) developing a SHAP explainer to automate feature selection and enhance model interpretability. Experiments show that Alioth achieves an average mean absolute error of 5.29% offline and 10.8% when testing on applications unseen in the training stage, outperforming the baseline methods. Alioth is also robust in signaling quality-of-service violation under dynamicity. Finally, we demonstrate a possible application of Alioth’s interpretability, providing insights to benefit the decision-making of cloud operators. The dataset and code of Alioth have been released on GitHub. Tianyao Shi, Yingxuan Yang, Yunlong Cheng, Xiaofeng Gao 0001, Yongqiang Yang |
IPDPS | 6 |
| 2023 | Prism: Revealing Hidden Functional Clusters from Massive Instances in Cloud SystemsabstractEnsuring the reliability of cloud systems is critical for both cloud vendors and customers. Cloud systems often rely on virtualization techniques to create instances of hardware resources, such as virtual machines. However, virtualization hinders the observability of cloud systems, making it challenging to diagnose platform-level issues. To improve system observability, we propose to infer functional clusters of instances, i.e., groups of instances having similar functionalities. We first conduct a pilot study on a large-scale cloud system, i.e., Huawei Cloud, demonstrating that instances having similar functionalities share similar communication and resource usage patterns. Motivated by these findings, we formulate the identification of functional clusters as a clustering problem and propose a non-intrusive solution called Prism. Prism adopts a coarse-to-fine clustering strategy. It first partitions instances into coarse-grained chunks based on communication patterns. Within each chunk, Prism further groups instances with similar resource usage patterns to produce fine-grained functional clusters. Such a design reduces noises in the data and allows Prism to process massive instances efficiently. We evaluate Prism on two datasets collected from the real-world production environment of Huawei Cloud. Our experiments show that Prism achieves a v-measure of ∼0.95, surpassing existing state-of-the-art solutions. Additionally, we illustrate the integration of Prism within monitoring systems for enhanced cloud reliability through two real-world use cases. Jinyang Liu 0002, Jiazhen Gu, Junjie Huang 0008, Zhuangbin Chen, Zengyin Yang, Yongqiang Yang, Michael R. Lyu |
ASE | 8 |
| 2023 | P4LRU: Towards An LRU Cache Entirely in Programmable Data PlaneabstractThe data plane cache, a critical functionality found in numerous network devices, such as programmable switches, intelligent NICs, and DPUs, is often subject to limitations in its programmability and memory access capacity. As a result, the majority of existing data plane caches rely on simple and inefficient replacement policies. This paper is set to introduce LRU, a near-optimal replacement policy, into the programmable data plane. We first explore the reasons why the traditional implementation of LRU is not suitable for deployment on the data plane. Consequently, we propose P4LRU, a pipeline-optimized version of the LRU implementation. Building on P4LRU, we conceive three distinct in-network systems - LruTable, LruIndex, and LruMon, and successfully bring them to life on Tofino switches. Our thorough experimental trials establish that P4LRU provides a significant performance boost over existing data plane caches in these three systems. We have open-sourced the source codes for the three systems on GitHub [1]. Yikai Zhao 0001, Wenrui Liu 0006, Fenghao Dong, Tong Yang 0003, Yuanpeng Li 0002, Kaicheng Yang 0001, Zirui Liu 0002, Zhengyi Jia, Yongqiang Yang |
SIGCOMM | 9 |
| 2023 | AAsclepius: Monitoring, Diagnosing, and Detouring at the Internet Peering Edge
Kaicheng Yang 0001, Yuanpeng Li 0002, Tong Yang 0003, Ruijie Miao, Yikai Zhao 0001, Chaoyang Ji, Penghui Mi, Qiong Xie, Hao Wang 0005, Yinhua Wang, Zhiqiang Liao, Chengqiang Huang, Yongqiang Yang |
USENIX ATC | 16 |
| 2022 | Adversarial Learning from CrowdsabstractLearning from Crowds (LFC) seeks to induce a high-quality classifier from training instances, which are linked to a range of possible noisy annotations from crowdsourcing workers under their various levels of skills and their own preconditions. Recent studies on LFC focus on designing new methods to improve the performance of the classifier trained from crowdsourced labeled data. To this day, however, there remain under-explored security aspects of LFC systems. In this work, we seek to bridge this gap. We first show that LFC models are vulnerable to adversarial examples---small changes to input data can cause classifiers to make prediction mistakes. Second, we propose an approach, A-LFC for training a robust classifier from crowdsourced labeled data. Our empirical results on three real-world datasets show that the proposed approach can substantially improve the performance of the trained classifier even with the existence of adversarial examples. On average, A-LFC has 10.05% and 11.34% higher test robustness than the state-of-the-art in the white-box and black-box attack settings, respectively. Hailong Sun 0001, Yongqiang Yang |
AAAI | 3 |
| 2022 | Characterizing and Mitigating Anti-patterns of Alerts in Industrial Cloud SystemsabstractAlerts are crucial for requesting prompt human intervention upon cloud anomalies. The quality of alerts significantly affects the cloud reliability and the cloud provider’s business revenue. In practice, we observe on-call engineers being hindered from quickly locating and fixing faulty cloud services because of the vast existence of misleading, non-informative, non-actionable alerts. We call the ineffectiveness of alerts "anti-patterns of alerts". To better understand the anti-patterns of alerts and provide actionable measures to mitigate anti-patterns, in this paper, we conduct the first empirical study on the practices of mitigating anti-patterns of alerts in an industrial cloud system. We study the alert strategies and the alert processing procedure at Huawei Cloud, a leading cloud provider. Our study combines the quantitative analysis of millions of alerts in two years and a survey with eighteen experienced engineers. As a result, we summarized four individual anti-patterns and two collective anti-patterns of alerts. We also summarize four current reactions to mitigate the anti-patterns of alerts, and the general preventative guidelines for the configuration of alert strategy. Lastly, we propose to explore the automatic evaluation of the Quality of Alerts (QoA), including the indicativeness, precision, and handleability of alerts, as a future research direction that assists in the automatic detection of alerts’ anti-patterns. The findings of our study are valuable for optimizing cloud monitoring systems and improving the reliability of cloud services. Jiacheng Shen, Yuxin Su 0001, Xiaoxue Ren, Yongqiang Yang, Michael R. Lyu |
DSN | 5 |
| 2022 | A Robust Service Mapping Scheme for Multi-Tenant CloudsabstractIn a multi-tenant cloud, cloud vendors provide services (e.g., elastic load-balancing, virtual private networks) on service nodes for tenants. Thus, the mapping of tenants’ traffic and service nodes is an important issue in multi-tenant clouds. In practice, unreliability of service nodes and uncertainty/dynamics of tenants’ traffic are two critical challenges that affect the tenants’ QoS. However, previous works often ignore the impact of these two challenges, leading to poor system robustness when encountering system accidents. To bridge the gap, this paper studies the problem of robust service mapping in multi-tenant clouds (RSMP). Due to traffic dynamics, we take a two-step approach:service node assignmentandtenant traffic scheduling. For service node assignment, we prove its NP-Hardness and analyze its problem difficulty. Then, we propose an efficient algorithm with bounded approximation factors based on randomized rounding and knapsack. For tenant traffic scheduling, we design an approximation algorithm based on fully polynomial time approximation scheme (FPTAS). The proposed algorithm achieves the approximation factor of 2+$\epsilon $, where$\epsilon $is an arbitrarily small value. Both small-scale experimental results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives. Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, Yutong Zhai, Qianyu Zhang 0001, He Huang 0001, Yongqiang Yang |
IEEE/ACM Trans. Netw. | 7 |
| 2021 | SketchINT: Empowering INT with TowerSketch for Per-flow Per-switch Measurementabstract1Network measurement is indispensable to network operations. Two most promising measurement solutions are In-band Network Telemetry (INT) solutions and sketching solutions. INT solutions provide fine-grained per-switch per-packet information at the cost of high network overhead. Sketching solutions have low network overhead but fail to achieve both simplicity and accuracy for per-flow measurement. To keep their advantages, and at the same time, overcome their shortcomings, we first design SketchINT to combine INT and sketches, aiming to obtain all per-flow per-switch information with low network overhead. Second, for deployment flexibility and measurement accuracy, we design a new sketch for SketchINT, namely TowerSketch, which achieves both simplicity and accuracy. The key idea of TowerSketch is to use different-sized counters for different arrays under the property that the number of bits used for different arrays stays the same. TowerSketch can automatically record larger flows in larger counters and smaller flows in smaller counters. We have fully implemented our SketchINT prototype on a testbed consisting of 10 switches. We also implement our TowerSketch on P4, single-core CPU, multi-core CPU, and FPGA platforms to verify its deployment flexibility. Extensive experimental results verify that 1) TowerSketch achieves better accuracy than prior art on various tasks, outperforming the state-of-the-art ElasticSketch up to 13.9 times in terms of error; 2) Compared to INT, SketchINT reduces the number of packets in the collection process by 3 4 orders of magnitude with an error smaller than 5%. Kaicheng Yang 0001, Yuanpeng Li 0002, Zirui Liu 0002, Tong Yang 0003, Yu Zhou 0008, Jintao He, Jing'an Xue, Zhengyi Jia, Yongqiang Yang |
ICNP | 10 |
| 2021 | Scalable On-Switch Rate Limiters for the CloudabstractWhile most clouds use on-server rate limiters for bandwidth allocation, we propose to implement them on switches. On-switch rate limiters can simplify network management and promote the performance of control-plane rate limiting applications. We leverage the recent progress of programmable switches to implement on-switch rate limiters, named SwRL. In the design of SwRL, we make design choices according to the programmable hardware characteristics, we deeply optimize the memory usage of the algorithm so as to fit a cloud-scale (one million) rate limiters in a single switch, and we complement the missing computation primitives of the hardware using a pre-computed approximate table. We further developed three control-plane applications and integrate them with SwRL, showing the control-plane interoperability of SwRL. We prototype and evaluate SwRL in both testbed and production environments, demonstrating its good properties of precision rate control, scalability, interoperability, and manageability (execution environmental isolation). Yongchao He, Wenfei Wu, Xuemin Wen, Yongqiang Yang |
INFOCOM | 5 |
| 2021 | Robust Service Mapping in Multi-Tenant CloudsabstractIn a multi-tenant cloud, cloud vendors provide services (e.g., elastic load-balancing, virtual private networks) on service nodes for tenants. Thus, the mapping of tenants' traffic and service nodes is an important issue in multi-tenant clouds. In practice, unreliability of service nodes and uncertainty/dynamics of tenants' traffic are two critical challenges that affect the tenants' QoS. However, previous works often ignore the impact of these two challenges, leading to poor system robustness when encountering system accidents. To bridge the gap, this paper studies the problem of robust service mapping in multi-tenant clouds (RSMP). Due to traffic dynamics, we take a two-step approach: service node assignment and tenant traffic scheduling. For service node assignment, we prove its NP-Hardness and analyze its problem difficulty. Then, we propose an efficient algorithm with bounded approximation factors based on randomized rounding and knapsack. For tenant traffic scheduling, we design an approximation algorithm based on fully polynomial time approximation scheme (FPTAS). The proposed algorithm achieves the approximation factor of 2+ ε , where ε is an arbitrarily small value. Both small-scale experimental results and large-scale simulation results show the superior performance of our proposed algorithms compared with other alternatives. Jingzhou Wang, Gongming Zhao, Hongli Xu 0001, He Huang 0001, Luyao Luo, Yongqiang Yang |
INFOCOM | 6 |
| 2021 | Graph-based Incident Aggregation for Large-Scale Online Service SystemsabstractAs online service systems continue to grow in terms of complexity and volume, how service incidents are managed will significantly impact company revenue and user trust. Due to the cascading effect, cloud failures often come with an overwhelming number of incidents from dependent services and devices. To pursue efficient incident management, related incidents should be quickly aggregated to narrow down the problem scope. To this end, in this paper, we propose GRLIA, an incident aggregation framework based on graph representation learning over the cascading graph of cloud failures. A representation vector is learned for each unique type of incident in an unsupervised and unified manner, which is able to simultaneously encode the topological and temporal correlations among incidents. Thus, it can be easily employed for online incident aggregation. In particular, to learn the correlations more accurately, we try to recover the complete scope of failures’ cascading impact by leveraging fine-grained system monitoring data, i.e., Key Performance Indicators (KPIs). The proposed framework is evaluated with real-world incident data collected from a large-scale online service system of Huawei Cloud. The experimental results demonstrate that GRLIA is effective and outperforms existing methods. Furthermore, our framework has been successfully deployed in industrial practice. Zhuangbin Chen, Jinyang Liu 0002, Yuxin Su 0001, Hongyu Zhang 0002, Xuemin Wen, Yongqiang Yang, Michael R. Lyu |
ASE | 7 |
| 2021 | AID: Efficient Prediction of Aggregated Intensity of Dependency in Large-scale Cloud SystemsabstractService reliability is one of the key challenges that cloud providers have to deal with. In cloud systems, unplanned service failures may cause severe cascading impacts on their dependent services, deteriorating customer satisfaction. Predicting the cascading impacts accurately and efficiently is critical to the operation and maintenance of cloud systems. Existing approaches identify whether one service depends on another via distributed tracing but no prior work focused on discriminating to what extent the dependency between cloud services is. In this paper, we survey the outages and the procedure for failure diagnosis in two cloud providers to motivate the definition of the intensity of dependency. We define the intensity of dependency between two services as how much the status of the callee service influences the caller service. Then we propose AID, the first approach to predict the intensity of dependencies between cloud services. AID first generates a set of candidate dependency pairs from the spans. AID then represents the status of each cloud service with a multivariate time series aggregated from the spans. With the representation of services, AID calculates the similarities between the statuses of the caller and the callee of each candidate pair. Finally, AID aggregates the similarities to produce a unified value as the intensity of the dependency. We evaluate AID on the data collected from an open-source microservice benchmark and a cloud system in production. The experimental results show that AID can efficiently and accurately predict the intensity of dependencies. We further demonstrate the usefulness of our method in a large-scale commercial cloud system. Jiacheng Shen, Yuxin Su 0001, Yongqiang Yang, Michael R. Lyu |
ASE | 5 |
| 2021 | Exploiting augmented intelligence in the modeling of safety-critical autonomous systemsabstractAbstract Machine learning (ML) is used increasingly in safety-critical systems to provide more complex autonomy to make the system to do decisions by itself in uncertain environments. Using ML to learn system features is fundamentally different from manually implementing them in conventional components written in source code. In this paper, we make a first step towards exploring the architecture modeling of safety-critical autonomous systems which are composed of conventional components and ML components, based on natural language requirements. Firstly, augmented intelligence for restricted natural language requirement modeling is proposed. In that, several AI technologies such as natural language processing and clustering are used to recommend candidate terms to the glossary, as well as machine learning is used to predict the category of requirements. The glossary including data dictionary and domain glossary and the category of requirements will be used in the restricted natural language requirement specification method RNLReq, which is equipped with a set of restriction rules and templates to structure and restrict the way how users document requirements. Secondly, automatic generation of SysML architecture models from the RNLReq requirement specifications is presented. Thirdly, the prototype tool is implemented based on Papyrus. Finally, it presents the evaluation of the proposed approach using an industrial autonomous guidance, navigation and control case study. Zhibin Yang 0005, Yang Bao 0007, Yongqiang Yang, Jean-Paul Bodeveix, Mamoun Filali, Zonghua Gu 0001 |
Formal Aspects Comput. | 3 |
| 2011 | JVM Virtual Method Invoking Optimization Based on CAM TableabstractIn Java programs, it needs to use the information of the method type to resolve the virtual method dynamically, which restricts the performance greatly. Currently, the solution is mainly the technique of inline caching, which can be divided into two categories: monomorphic inline caching and polymorphic inline caching. Because of the simple implementation of monomorphic inline caching, it is more commonly used, but it cannot resolve the problem of frequently seeing different types of objects at one call site. Although polymorphic inline caching can solve this problem, it is costly. This paper presents the CAM hardware table, proposes a virtual method invoking mechanism based on software and hardware co-design. It solves the problem of frequently seeing different types of objects at one call site, and does not introduce additional overhead. The experimental evaluation shows that it improves the cached hit rate from 13.3% to 76.4% and improves 16.2% performance of the virtual test, it improves the performance of SPECjvm 98 by 6.4% on average. Songsong Cai, Yongqiang Yang, Chuanwen Lin |
NAS | 2 |