VLDB 2026 Research / reviewers in the wild / expert
Si Qin
dblp:148/9596
· DBLP profile ↗
33ranked-venue papers
4as first author
24since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 2 since 2021Databases, data management, data science and information retrieval · 8 · 8 since 2021Software engineering, systems software and programming languages · 7 · 6 since 2021Systems, architecture and hardware · 4 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Medical multi-recall embedding: Adaptive retrieval for diverse evidence in medical RAG systems
Changjin Li, Fengshi Jing, Huarun Li, Zhougzhi Xu, Huiru Zou, Qiting Wang, Yuchen Qian, Boyu Cao, Si Qin, Weibin Cheng, Haobin Zhang |
Inf. Process. Manag. | 12 |
| 2026 | TDBCL: A time series dual-branch balance contrastive learning for imbalanced classification
Haobin Zhang, Shengning Chan, Zhongzhi Xu, Fengshi Jing, Huiru Zou, Si Qin, Weibin Cheng |
Pattern Recognit. | 8 |
| 2025 | AllHands :Ask Me Anything on Large-scale Verbatim Feedback via Large Language ModelsabstractVerbatim feedback constitutes a valuable repository of user experiences, opinions, and requirements, crucial for data engineering and software development. Extracting meaningful insights from large-scale feedback data presents a significant challenge. This paper introduces Allhands, an innovative ana-lytic framework that transforms traditional large-scale feedback analysis tasks through a natural language interface, leveraging large language models (LLMs). Allhands performs initial classification and topic modeling on feedback to convert it into a structurally augmented format, enhancing accuracy, robustness and generalization with the aid of LLMs. Subsequently, an LLM-based code-first agent interprets users' diverse natural language questions about the feedback, automatically translates them into executable call of analytic tools or code, and delivers comprehensive multi-modal responses, including text, code, tables, and images. This eliminates the need for developing individual feedback analytic tools for each request, reducing human effort and making the system more accessible and flexible to users. We evaluate Allhands across three diverse feedback datasets, demonstrating its superior efficacy in all stages of analysis, from classification and topic modeling to providing an “ask me anything” experience with comprehensive, accurate, and human-readable responses. To the best of our knowl-edge, Allhands is the first comprehensive feedback analysis framework supporting diverse and customized insight extraction requirements through a natural language interface. Chaoyun Zhang, Zicheng Ma, Shilin He, Si Qin, Minghua Ma, Xiaoting Qin, Yu Kang 0006, Yuyi Liang, Xiaoyu Gou, Yajie Xue, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
ICDE | 5 |
| 2025 | RuAG: Learned-rule-augmented Generation for Large Language ModelsabstractIn-context learning (ICL) and Retrieval-Augmented Generation (RAG) have gained attention for their ability to enhance LLMs' reasoning by incorporating external knowledge but suffer from limited contextual window size, leading to insufficient information injection. To this end, we propose a novel framework to automatically distill large volumes of offline data into interpretable first-order logic rules, which are injected into LLMs to boost their reasoning capabilities. Our method begins by formulating the search process relying on LLMs' commonsense, where LLMs automatically define head and body predicates. Then, we apply Monte Carlo Tree Search (MCTS) to address the combinational searching space and efficiently discover logic rules from data. The resulting logic rules are translated into natural language, allowing targeted knowledge injection and seamless integration into LLM prompts for LLM's downstream task reasoning. We evaluate our framework on public and private industrial tasks, including Natural Language Processing (NLP), time-series, decision-making, and industrial tasks, demonstrating its effectiveness in enhancing LLM's capability over diverse tasks. Yudi Zhang 0006, Pei Xiao 0005, Lu Wang 0029, Chaoyun Zhang, Yali Du 0001, Yevgeniy Puzyrev, Randolph Yao, Si Qin, Qingwei Lin, Mykola Pechenizkiy, Dongmei Zhang 0001, Saravan Rajmohan, Qi Zhang 0066 |
ICLR | 9 |
| 2025 | Label Distribution Learning with Biased Annotations Assisted by Multi-Label LearningabstractMulti-label learning (MLL) has gained attention for its ability to represent real-world data. Label Distribution Learning (LDL), an extension of MLL to learning from label distributions, faces challenges in collecting accurate label distributions. To address the issue of biased annotations, based on the low-rank assumption, existing works recover true distributions from biased observations by exploring the label correlations. However, recent evidence shows that the label distribution tends to be full-rank, and naive apply of low-rank approximation on biased observation leads to inaccurate recovery and performance degradation. In this paper, we address the LDL with biased annotations problem from a novel perspective, where we first degenerate the soft label distribution into a hard multi-hot label and then recover the true label information for each instance. This idea stems from an insight that assigning hard multi-hot labels is often easier than assigning a soft label distribution, and it shows stronger immunity to noise disturbances, leading to smaller label bias. Moreover, assuming that the multi-label space for predicting label distributions is low-rank offers a more reasonable approach to capturing label correlations. Theoretical analysis and experiments confirm the effectiveness and robustness of our method on real-world datasets. Zhiqiang Kou, Si Qin, Hailin Wang 0001, Jing Wang 0113, Ming-Kun Xie, Shuo Chen 0003, Yuheng Jia, Tongliang Liu, Masashi Sugiyama, Xin Geng 0001 |
IJCAI | 2 |
| 2025 | Large Language Models can Deliver Accurate and Interpretable Time Series Anomaly DetectionabstractTime series anomaly detection (TSAD) plays a crucial role in various industrial applications. Traditional deep learning TSAD models require extensive training data and operate as black boxes, lacking interpretability for detected anomalies. To address these challenges, we propose LLMAD, a novel TSAD method that employs Large Language Models (LLMs) to deliver accurate and interpretable TSAD results. LLMAD applies in-context anomaly detection by retrieving both positive and negative similar time series segments, significantly enhancing LLMs' effectiveness. Furthermore, LLMAD employs the Anomaly Detection Chain-of-Thought approach to mimic expert logic for its decision-making process. This further enhances its performance and enables LLMAD to provide explanations for their detections through versatile perspectives. Chaoyun Zhang, Jiaxu Qian, Minghua Ma, Si Qin, Chetan Bansal, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001 |
KDD (2) | 5 |
| 2025 | UFO: A UI-Focused Agent for Windows OS InteractionabstractChaoyun Zhang, Liqun Li, Shilin He, Xu Zhang, Bo Qiao, Si Qin, Minghua Ma, Yu Kang, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang, Qi Zhang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Chaoyun Zhang, Liqun Li, Shilin He, Xu Zhang 0024, Bo Qiao 0001, Si Qin, Minghua Ma, Yu Kang 0006, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang 0001, Qi Zhang 0066 |
NAACL (Long Papers) | 6 |
| 2025 | GUI-Actor: Coordinate-Free Visual Grounding for GUI AgentsabstractOne of the principal challenges in building VLM-powered GUI agents is visual grounding—localizing the appropriate screen region for action execution based on both the visual content and the textual plans. Most existing work formulates this as a text-based coordinate generation task. However, these approaches suffer from several limitations: weak spatial-semantic alignment due to lack of explicit spatial supervision; inability to handle ambiguous supervision targets, as single-point predictions penalize valid variations; and a mismatch between the dense nature of screen coordinates and the coarse, patch-level granularity of visual features extracted by models like Vision Transformers. In this paper, we propose **GUI-Actor**, a VLM-based method for coordinate-free GUI grounding. At its core, **GUI-Actor** introduces an attention-based action head that learns to align a dedicated `<ACTOR>` token with all relevant visual patch tokens, enabling the model to propose one or more action regions in a single forward pass. In line with this, we further design a grounding verifier to evaluate and select the most plausible action region from the candidates proposed for action execution. Extensive experiments show that **GUI-Actor** outperforms prior state-of-the-art methods on multiple GUI action grounding benchmarks, with improved generalization to unseen screen resolutions and layouts. Notably, **GUI-Actor-7B** achieves scores of **40.7** with Qwen2-VL and **44.6** with Qwen2.5-VL as backbones, outperforming **UI-TARS-72B (38.1)** on ScreenSpot-Pro, with significantly fewer parameters and training data. Furthermore, by incorporating the verifier, we find that fine-tuning only the newly introduced action head (~100M parameters for 7B model) while keeping the VLM backbone frozen is sufficient to achieve performance comparable to previous state-of-the-art models, highlighting that **GUI-Actor** can endow the underlying VLM with effective grounding capabilities without compromising its general-purpose strengths. Project page: [https://aka.ms/GUI-Actor](https://aka.ms/GUI-Actor) Qianhui Wu, Kanzhi Cheng, Rui Yang 0010, Chaoyun Zhang, Huiqiang Jiang, Jian Mu, Baolin Peng, Bo Qiao 0001, Reuben Tan, Si Qin, Lars Liden, Qingwei Lin, Huan Zhang 0001, Tong Zhang 0001, Dongmei Zhang 0001, Jianfeng Gao 0001 |
NeurIPS | 11 |
| 2024 | COIN: Chance-Constrained Imitation Learning for Safe and Adaptive Resource Oversubscription under UncertaintyabstractWe address the real problem of safe, robust, adaptive resource oversubscription in uncertain environments with our proposed novel technique of chance-constrained imitation learning. Our objective is to enhance resource efficiency while ensuring safety against congestion risk. Traditional supervised or forecasting models are ineffective in learning adaptive oversubscription policies, and conventional online optimization or reinforcement learning is difficult to deploy on real systems. Offline policy learning methods, such as Imitation Learning (IL) can leverage historical resource utilization telemetry data to learn effective policies if we can ensure robustness and safety from the underlying uncertainty in the domain, and thus the data. Our work investigates the nature of this uncertainty, how it can be quantified and proposes a novel chance-constrained IL that implicitly models such uncertainty in a principled manner via additional knowledge in the form of stochastic constraints on the associated risk, to learn provably safe and robust policies. We show empirically a substantial improvement (~ 3-4×) in capacity efficiency and congestion safety in test as well as real deployments. Lu Wang 0029, Mayukh Das, Fangkai Yang, Bo Qiao 0001, Hang Dong 0004, Chetan Bansal, Si Qin, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001, Qi Zhang 0066 |
CIKM | 8 |
| 2024 | Deoxys: A Causal Inference Engine for Unhealthy Node Mitigation in Large-scale Cloud InfrastructureabstractThe presence of unhealthy nodes in cloud infrastructure signals the potential failure of machines, which can significantly impact the availability and reliability of cloud services, resulting in negative customer experiences. Effectively addressing unhealthy node mitigation is therefore vital for sustaining cloud system performance. This paper introduces Deoxys, a causal inference engine tailored to recommending mitigation actions for unhealthy node in cloud systems to minimize virtual machine downtime and interruptions during unhealthy events. It employs double machine learning combined with causal forest to produce precise and reliable mitigation recommendations based solely on limited observational data collected from the historical unhealthy events. To enhance the causal inference model, Deoxys further incorporates a policy fallback mechanism based on model uncertainty and action overriding mechanisms to (i) improve the reliability of the system, and (ii) strike a good tradeoff between downtime reduction and resource utilization, thereby enhancing the overall system performance. Chaoyun Zhang, Randolph Yao, Si Qin, Ze Li 0005, Shekhar Agrawal, Binit R. Mishra, Minghua Ma, Qingwei Lin, Murali Chintalapati, Dongmei Zhang 0001 |
SoCC | 3 |
| 2024 | Xpert: Empowering Incident Management with Query Recommendations via Large Language ModelsabstractLarge-scale cloud systems play a pivotal role in modern IT infrastructure. However, incidents occurring within these systems can lead to service disruptions and adversely affect user experience. To swiftly resolve such incidents, on-call engineers depend on crafting domain-specific language (DSL) queries to analyze telemetry data. However, writing these queries can be challenging and time-consuming. This paper presents a thorough empirical study on the utilization of queries of KQL, a DSL employed for incident management in a large-scale cloud management system at Microsoft. The findings obtained underscore the importance and viability of KQL queries recommendation to enhance incident management. Yuxuan Jiang 0016, Chaoyun Zhang, Shilin He, Minghua Ma, Si Qin, Yu Kang 0006, Yingnong Dang, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
ICSE | 6 |
| 2024 | SMDE: Unsupervised representation learning for time series based on signal mode decomposition and ensemble
Haobin Zhang, Shengning Chan, Si Qin |
Knowl. Based Syst. | 3 |
| 2024 | Counter-Empirical Attacking Based on Adversarial Reinforcement Learning for Time-Relevant Scoring SystemabstractScoring systems are commonly seen for platforms in the era of Big Data. From credit scoring systems in financial services to membership scores in E-commerce shopping platforms, platform managers use such systems to guide users towards the encouraged activity pattern, and manage resources more effectively and efficiently. To establish such scoring systems, several “empirical criteria” are first determined, followed by a dedicated top-down design for each score factor, which usually requires enormous effort to adjust and tune the scoring function in the new application scenario. What's worse, many fresh projects usually have no ground truth or any experience to evaluate a reasonable scoring system, making the designing even harder. To reduce the effort of manual adjustment of the scoring function in every new scoring system, we innovatively study the scoring system from the preset empirical criteria without any ground truth and propose a novel framework to improve the system from scratch. In this paper, we propose a “counter-empirical attacking” mechanism that can generate “attacking” behavior traces and try to break the empirical rules of the scoring system. Then an adversarial “enhancer” is applied to evaluate the scoring system and find the improvement strategy. By training the adversarial learning problem, a proper scoring function can be learned to be robust to the attacking activity traces that are trying to violate the empirical criteria. Extensive experiments have been conducted on two scoring systems, including a shared computing resource platform and a financial credit system. The experimental results have validated the effectiveness of our proposed framework. Xiangguo Sun, Hong Cheng 0001, Hang Dong 0004, Bo Qiao 0001, Si Qin, Qingwei Lin |
IEEE Trans. Knowl. Data Eng. | 5 |
| 2023 | Snape: Reliable and Low-Cost Computing with Mixture of Spot and On-Demand VMsabstractCloud providers often have resources that are not being fully utilized, and they may offer them at a lower cost to make up for the reduced availability of these resources. However, customers may be hesitant to use such offerings (such as spot VMs) as making trade-offs between cost and resource availability is not always straightforward. In this work, we propose Snape (Spot On-demand Perfect Mixture), an intelligent framework to optimize the cost and resource availability by dynamically mixing on-demand VMs with spot VMs. Through a detailed characterization based on real production traces, we verify that the eviction of spot VMs is predictable to some extent. Snape also leverages constrained reinforcement learning to adjust the mixture policy online. Experiments across different configurations show that Snape achieves 44% savings compared to using only on-demand VMs while maintaining 99.96% availability, which is 2.77% higher than using only spot VMs. Fangkai Yang, Lu Wang 0029, Zhenyu Xu 0003, Liqun Li, Bo Qiao 0001, Camille Couturier, Chetan Bansal, Soumya Ram, Si Qin, Íñigo Goiri, Eli Cortez, Terry Yang, Victor Rühle, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
ASPLOS (3) | 10 |
| 2023 | How Different are the Cloud Workloads? Characterizing Large-Scale Private and Public Cloud WorkloadsabstractWith the rapid development of cloud systems, an increasing number of service workloads are deployed in the private cloud and/or public cloud. Although large cloud providers such as Azure and Google have published workload traces in the past, prior work has not focused on analyzing and characterizing the differences between private and public cloud workloads in detail. Based on our experience working with Azure, one of the most widely used cloud platforms in the world, we find that the workload characteristics are different between the private and public cloud workloads. Specifically, compared with the public cloud workloads, the private cloud workloads tend to be more homogeneous in both deployment sizes and utilization patterns, more static with occasional bursts in deployment characteristics, and more region-agnostic regarding the sensitivity to deployed regions. Our findings gain several insights and implications on cloud management and motivate us to build a centralized workload knowledge base. Xiaoting Qin, Minghua Ma, Yuheng Zhao, Anjaly Parayil, Chetan Bansal, Saravan Rajmohan, Íñigo Goiri, Eli Cortez, Si Qin, Qingwei Lin, Dongmei Zhang 0001 |
DSN | 12 |
| 2023 | CODEC: Cost-Effective Duration Prediction System for Deadline Scheduling in the CloudabstractModern cloud platforms allow customers to flexibly allocate or release computing resources. One crucial scenario is how to drive existing VMs to a specific state by a given deadline in a reliable and cost-effective manner. These state transition requests could involve starting or reallocating VMs. Performing millions of these requests per day can be challenging because the throughput of cloud platforms is not consistently deterministic. To meet customer service level agreements, cloud providers often trade-off cost for reliability and use oversized estimates to ensure that they initiate state transitions against VMs well ahead of time. In this paper, we propose a COst-effective Duration prEdiCtion system (CODEC) that solves the deadline scheduling problem for cloud providers by building intelligent automation to discover and execute optimal strategies when performing concurrent requests. In the CODEC, the core is to categorize durations of requests into buckets in real-time and the buffer time of each bucket is then predicted by extreme value theory. Extensive experiments show the proposed approach can guarantee the success rate of requests, while significantly saving cost with 38.71% and 86.97% compared to the baseline approach and the static buffer time, respectively. The results obtained show that CODEC can effectively and efficiently solve the deadline scheduling problem, and lead a worldwide cloud provider to deploy its prototype. Minghua Ma, Si Qin, Bo Qiao 0001, Randolph Yao, Harshwardhan Chaturvedi, Murali Chintalapati, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
ISSRE | 4 |
| 2023 | Learning Cooperative Oversubscription for Cloud by Chance-Constrained Multi-Agent Reinforcement LearningabstractOversubscription is a common practice for improving cloud resource utilization. It allows the cloud service provider to sell more resources than the physical limit, assuming not all users would fully utilize the resources simultaneously. However, how to design an oversubscription policy that improves utilization while satisfying some safety constraints remains an open problem. Existing methods and industrial practices are over-conservative, ignoring the coordination of diverse resource usage patterns and probabilistic constraints. To address these two limitations, this paper formulates the oversubscription for cloud as a chance-constrained optimization problem and proposes an effective Chance-Constrained Multi-Agent Reinforcement Learning (C2MARL) method to solve this problem. Specifically, C2MARL reduces the number of constraints by considering their upper bounds and leverages a multi-agent reinforcement learning paradigm to learn a safe and optimal coordination policy. We evaluate our C2MARL on an internal cloud platform and public cloud datasets. Experiments show that our C2MARL outperforms existing methods in improving utilization () under different levels of safety constraints. Junjie Sheng, Lu Wang 0029, Fangkai Yang, Bo Qiao 0001, Hang Dong 0004, Xiangfeng Wang 0001, Bo Jin 0003, Jun Wang 0006, Si Qin, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001 |
WWW | 9 |
| 2023 | Multi-Task Learning With Hierarchical Guidance for Locating and Stratifying Submucosal TumorsabstractLocating and stratifying the submucosal tumor of the digestive tract from endoscopy ultrasound (EUS) images are of vital significance to the preliminary diagnosis of tumors. However, the above problems are challenging, due to the poor appearance contrast between different layers of the digestive tract wall (DTW) and the narrowness of each layer. Few of existing deep-learning based diagnosis algorithms are devised to tackle this issue. In this article, we build a multi-task framework for simultaneously locating and stratifying the submucosal tumor. And considering the awareness of the DTW is critical to the localization and stratification of the tumor, we integrate the DTW segmentation task into the proposed multi-task framework. Except for sharing a common backbone model, the three tasks are explicitly directed with a hierarchical guidance module, in which the probability map of DTW itself is used to locally enhance the feature representation for tumor localization, and the probability maps of DTW and tumor are jointly employed to locally enhance the feature representation for tumor stratification. Moreover, by means of the dynamic class activation map, probability maps of DTW and tumor are reused to enforce the stratification inference process to pay more attention to DTW and tumor regions, contributing to a reliable and interpretable submucosal tumor stratification model. Additionally, considering the relation with respect to other structures is beneficial for stratifying tumors, we devise a graph reasoning module to replenish non-local relation knowledge for the stratification branch. Experiments on a Stomach-Esophagus and an Intestinal EUS dataset prove that our method achieves very appealing performance on both tumor localization and stratification, significantly outperforming state-of-the-art object detection approaches. Ruifei Zhang, Si Qin, Chaowei Fang, Guanbin Li, Xutao Lin |
IEEE J. Biomed. Health Informatics | 3 |
| 2022 | Solving the Batch Stochastic Bin Packing Problem in Cloud: A Chance-constrained Optimization ApproachabstractThis paper investigates a critical resource allocation problem in the first party cloud: scheduling containers to machines. There are tens of services, and each service runs a set of homogeneous containers with dynamic resource usage; containers of a service are scheduled daily in a batch fashion. This problem can be naturally formulated as Stochastic Bin Packing Problem (SBPP). However, traditional SBPP research often focuses on cases of empty machines, whose objective, i.e., to minimize the number of used machines, is not well-defined for the more common reality with nonempty machines. This paper aims to close this gap. First, we define a new objective metric, Used Capacity at Confidence (UCaC), which measures the maximum used resources at a probability and is proved to be consistent for both empty and nonempty machines and reformulate the SBPP under chance constraints. Second, by modeling the container resource usage distribution in a generative approach, we reveal that UCaC can be approximated with Gaussian, which is verified by trace data of real-world applications. Third, we propose an exact solver by solving the equivalent cutting stock variant as well as two heuristics-based solvers -- UCaC best fit, bi-level heuristics. We experimentally evaluate these solvers on both synthetic datasets and real application traces, demonstrating our methodology's advantage over traditional SBPP optimal solver minimizing the number of used machines, with a low rate of resource violations. Yunlei Lu, Liting Chen, Si Qin, Yixin Fang, Qingwei Lin, Thomas Moscibroda, Saravan Rajmohan, Dongmei Zhang 0001 |
KDD | 4 |
| 2022 | RESIN: A Holistic Service for Dealing with Memory Leaks in Production Cloud Infrastructure
Chang Lou, Peng Huang 0005, Yingnong Dang, Si Qin, Xinsheng Yang, Xukun Li, Qingwei Lin, Murali Chintalapati |
OSDI | 5 |
| 2021 | Predictive Job Scheduling under Uncertain Constraints in Cloud ComputingabstractCapacity management has always been a great challenge for cloud platforms due to massive, heterogeneous on-demand instances running at different times. To better plan the capacity for the whole platform, a class of cloud computing instances have been released to collect computing demands beforehand. To use such instances, users are allowed to submit jobs to run for a pre-specified uninterrupted duration in a flexible range of time in the future with a discount compared to the normal on-demand instances. Proactively scheduling those pre-collected job requests considering the capacity status over the platform can greatly help balance the computing workloads along time. In this work, we formulate the scheduling problem for these pre-collected job requests under uncertain available capacity as a Prediction + Optimization problem with uncertainty in constraints, and propose an effective algorithm called Controlling under Uncertain Constraints (CUC), where the predicted capacity guides the optimization of job scheduling and job scheduling results are leveraged to improve the prediction of capacity through Bayesian optimization. The proposed formulation and solution are commonly applicable for proactively scheduling problems in cloud computing. Our extensive experiments on three public, industrial datasets shows that CUC has great potential for supporting high reliability in cloud platforms. Hang Dong 0004, Boshi Wang, Bo Qiao 0001, Wenqian Xing, Chuan Luo 0002, Si Qin, Qingwei Lin, Dongmei Zhang 0001, Gurpreet Virdi, Thomas Moscibroda |
IJCAI | 6 |
| 2021 | HALO: Hierarchy-aware Fault Localization for Cloud SystemsabstractA typical cloud system has a large amount of telemetry data collected by pervasive software monitors that keep tracking the health status of the system. The telemetry data is essentially multi-dimensional data, which contains attributes and failure/success status of the system being monitored. By identifying the attribute value combinations where the failures are mostly concentrated (which we call fault-indicating combination), we can localize the cause of system failures into a smaller scope, thus facilitating fault diagnosis. However, due to the combinatorial explosion problem and the latent hierarchical structure in cloud telemetry data, it is still intractable to localize the fault to a proper granularity in an efficient way. In this paper, we propose HALO, a hierarchy-aware fault localization approach for locating the fault-indicating combinations from telemetry data. Our approach automatically learns the hierarchical relationship among attributes and leverages the hierarchy structure for precise and efficient fault localization. We have evaluated HALO on both industrial and synthetic datasets and the results confirm that HALO outperforms the existing methods. Furthermore, we have successfully deployed HALO to different services in Microsoft Azure and Microsoft 365, witnessed its impact in real-world practice. Xu Zhang 0024, Yong Xu 0010, Hongyu Zhang 0002, Si Qin, Ze Li 0005, Qingwei Lin, Yingnong Dang, Andrew Zhou, Saravanakumar Rajmohan, Dongmei Zhang 0001 |
KDD | 6 |
| 2021 | Effective low capacity status prediction for cloud systemsabstractIn cloud systems, an accurate capacity planning is very important for cloud provider to improve service availability. Traditional methods simply predicting "when the available resources is exhausted" are not effective due to customer demand fragmentation and platform allocation constraints. In this paper, we propose a novel prediction approach which proactively predicts the level of resource allocation failures from the perspective of low capacity status. By jointly considering the data from different sources in both time series form and static form, the proposed approach can make accurate LCS predictions in a complex and dynamic cloud environment, and thereby improve the service availability of cloud systems. The proposed approach is evaluated by real-world datasets collected from a large scale public cloud platform, and the results confirm its effectiveness. Hang Dong 0004, Si Qin, Yong Xu 0010, Bo Qiao 0001, Shandan Zhou, Xian Yang 0001, Chuan Luo 0002, Pu Zhao 0004, Qingwei Lin, Hongyu Zhang 0002, Abulikemu Abuduweili, Sanjay Ramanujan, Karthikeyan Subramanian, Andrew Zhou, Saravanakumar Rajmohan, Dongmei Zhang 0001, Thomas Moscibroda |
ESEC/SIGSOFT FSE | 2 |
| 2021 | Onion: identifying incident-indicating logs for cloud systemsabstractIn cloud systems, incidents affect the availability of services and require quick mitigation actions. Once an incident occurs, operators and developers often examine logs to perform fault diagnosis. However, the large volume of diverse logs and the overwhelming details in log data make the manual diagnosis process time-consuming and error-prone. In this paper, we propose Onion, an automatic solution for precisely and efficiently locating incident-indicating logs, which can provide useful clues for diagnosing the incidents. We first point out three criteria for localizing incident-indicating logs, i.e., Consistency, Impact, and Bilateral-Difference. Then we propose a novel agglomeration of logs, called log clique, based on which these criteria are satisfied. To obtain log cliques, we develop an incident-aware log representation and a progressive log clustering technique. Contrast analysis is then performed on the cliques to identify the incident-indicating logs. We have evaluated Onion using well-labeled log datasets. Onion achieves an average F1-score of 0.95 and can process millions of logs in only a few minutes, demonstrating its effectiveness and efficiency. Onion has also been successfully applied to the cloud system of Microsoft. Its practicability has been confirmed through the quantitative and qualitative analysis of the real incident cases. Xu Zhang 0024, Yong Xu 0010, Si Qin, Shilin He, Bo Qiao 0001, Ze Li 0005, Hongyu Zhang 0002, Xukun Li, Yingnong Dang, Qingwei Lin, Murali Chintalapati, Saravanakumar Rajmohan, Dongmei Zhang 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2020 | Efficient incident identification from multi-dimensional issue reports via meta-heuristic searchabstractIn large-scale cloud systems, unplanned service interruptions and outages may cause severe degradation of service availability. Such incidents can occur in a bursty manner, which will deteriorate user satisfaction. Identifying incidents rapidly and accurately is critical to the operation and maintenance of a cloud system. In industrial practice, incidents are typically detected through analyzing the issue reports, which are generated over time by monitoring cloud services. Identifying incidents in a large number of issue reports is quite challenging. An issue report is typically multi-dimensional: it has many categorical attributes. It is difficult to identify a specific attribute combination that indicates an incident. Existing methods generally rely on pruning-based search, which is time-consuming given high-dimensional data, thus not practical to incident detection in large-scale cloud systems. In this paper, we propose MID (Multi-dimensional Incident Detection), a novel framework for identifying incidents from large-amount, multi-dimensional issue reports effectively and efficiently. Key to the MID design is encoding the problem into a combinatorial optimization problem. Then a specific-tailored meta-heuristic search method is designed, which can rapidly identify attribute combinations that indicate incidents. We evaluate MID with extensive experiments using both synthetic data and real-world data collected from a large-scale production cloud system. The experimental results show that MID significantly outperforms the current state-of-the-art methods in terms of effectiveness and efficiency. Additionally, MID has been successfully applied to Microsoft's cloud systems and helped greatly reduce manual maintenance effort. Jiazhen Gu, Chuan Luo 0002, Si Qin, Bo Qiao 0001, Qingwei Lin, Hongyu Zhang 0002, Ze Li 0005, Yingnong Dang, Shaowei Cai 0001, Wei Wu 0011, Yangfan Zhou 0002, Murali Chintalapati, Dongmei Zhang 0001 |
ESEC/SIGSOFT FSE | 3 |
| 2020 | Improved two-dimensional DOA estimation using parallel coprime arrays
Si Qin, Yimin Zhang 0001, Moeness G. Amin |
Signal Process. | 1 |
| 2020 | CloudDet: Interactive Visual Analysis of Anomalous Performances in Cloud Computing SystemsabstractDetecting and analyzing potential anomalous performances in cloud computing systems is essential for avoiding losses to customers and ensuring the efficient operation of the systems. To this end, a variety of automated techniques have been developed to identify anomalies in cloud computing. These techniques are usually adopted to track the performance metrics of the system (e.g., CPU, memory, and disk I/O), represented by a multivariate time series. However, given the complex characteristics of cloud computing data, the effectiveness of these automated methods is affected. Thus, substantial human judgment on the automated analysis results is required for anomaly interpretation. In this paper, we present a unified visual analytics system named CloudDet to interactively detect, inspect, and diagnose anomalies in cloud computing systems. A novel unsupervised anomaly detection algorithm is developed to identify anomalies based on the specific temporal patterns of the given metrics data (e.g., the periodic pattern). Rich visualization and interaction designs are used to help understand the anomalies in the spatial and temporal context. We demonstrate the effectiveness of CloudDet through a quantitative evaluation, two case studies with real-world data, and interviews with domain experts. Yun Wang 0012, Leni Yang, Yifang Wang 0001, Bo Qiao 0001, Si Qin, Yong Xu 0010, Huamin Qu |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2019 | Cross-dataset Time Series Anomaly Detection for Cloud Systems
Xu Zhang 0024, Qingwei Lin, Yong Xu 0010, Si Qin, Hongyu Zhang 0002, Bo Qiao 0001, Yingnong Dang, Xinsheng Yang, Murali Chintalapati, Youjiang Wu, Ken Hsieh, Kaixin Sui, Yaohai Xu, Wenchi Zhang, Furao Shen, Dongmei Zhang 0001 |
USENIX ATC | 4 |
| 2017 | Robust DOA estimation in the presence of mis-calibrated sensorsabstractIn this paper, we consider robust direction-of-arrival (DOA) estimation for an array that contains mis-calibrated sensors with unknown gain and phase uncertainties. We develop two robust DOA estimation algorithms based on the maximum correntropy criterion (MCC). In the first algorithm, adaptively optimized weighting factors are obtained and applied to each sensor to effectively mitigate the effect of calibration error and array manifold distortions, and the results are fed into sparse reconstruction methods for DOA estimation. In the second algorithm, we further estimate the gain and phase errors of the mis-calibrated sensors so that the entire array is fully calibrated for improved DOA estimation. The effectiveness of the proposed techniques is verified using simulation results. Ben Wang 0002, Si Qin, Yimin Zhang 0001, Moeness G. Amin |
ICASSP | 2 |
| 2017 | DOA estimation exploiting a uniform linear array with multiple co-prime frequencies
Si Qin, Yimin Zhang 0001, Moeness G. Amin, Braham Himed |
Signal Process. | 1 |
| 2016 | Generalized coprime sampling of Toeplitz matricesabstractIncreased demand on spectrum sensing over a broad frequency band requires a high sampling rate and thus leads to a prohibitive volume of data samples. In some applications, e.g., spectrum estimation, only the second-order statistics are required. In this case, we may use a reduced data sampling rate by exploiting a low-dimensional representation of the original high-dimensional signals. In particular, the covariance matrix can be reconstructed from compressed data by utilizing its specific structure, e.g., the Toeplitz property. In this paper, we propose a general coprime sampling concept that implements effective compression of Toeplitz covariance matrices. Given a fixed number of data samples, we examine different schemes on covariance matrix acquisition, based on segmented data sequences. The effectiveness of the proposed technique is verified using simulation results. Si Qin, Yimin Zhang 0001, Moeness G. Amin, Abdelhak M. Zoubir |
ICASSP | 1 |
| 2015 | Doa estimation of nonparametric spreading spatial spectrum based on bayesian compressive sensing exploiting intra-task dependencyabstractFor spatially distributed targets encountered in radar and sonar applications, direct application of subspace-based methods usually do not lead to an accurate estimation of the direction and angular extent of the signal arrivals. If the spatial distribution of the targets can be parameterized with a known model a priori, the direction-of-arrival (DOA) estimation problems can be simplified as parameter estimation problems. However, these methods do not apply when the targets are not parameterizable. Motivated by this fact, we propose an effective approach for the DOA estimation of nonparametric spatially extended targets. In the proposed approach, the spatially extended targets are modeled as a continuous sparse structure, which are effectively estimated using the Bayesian compressive sensing techniques based on a paired spike-and-slab prior accounting for the angular target spread. In particular, the problem is examined under a collocated multiple-input multiple-output (MIMO) radar platform. Signal transmission at multiple coprime transmit frequencies are also considered to achieve increased degrees-of-freedom. The group sparsity of the targets across different frequencies is exploited to achieve improved DOA estimation performance. Si Qin, Qisong Wu, Yimin Zhang 0001, Moeness G. Amin |
ICASSP | 1 |
| 2014 | Doa estimation exploiting coprime arrays with sparse sensor spacingabstractIn this paper, we propose effective coprime array configurations in which the minimum interelement spacing is much larger than the typical half-wavelength requirement. Such configurations are important in many applications where the half-wavelength requirement cannot be met due to the physical sensors size or to avoid spatial oversampling in wideband operations. The application of such coprime arrays in direction-of-arrival estimations is examined using different algorithms. Yimin Zhang 0001, Si Qin, Moeness G. Amin |
ICASSP | 2 |