Zhe Xie

dblp:72/9947 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0001-9749-2539ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 TechSupportEval: An Automated Evaluation Framework for Technical Support Question Answering
abstract
Technical support question-answering (QA) systems assist users in diagnosing and resolving technical issues, but ensuring their reliability remains a challenge. Existing QA systems may generate inaccurate responses due to LLM hallucinations and retrieval errors, which can lead to misleading guidance. A reliable evaluation framework is essential for systematically improving technical support QA systems, ensuring they generate accurate guidance. However, existing evaluation methods for QA systems struggle to precisely match key terms, and verify step order and completeness.To address these challenges, we propose TechSupportEval, an automated evaluation framework for technical support QA. Our framework introduces two novel techniques: (1) ClozeFact, which formulates fact verification as a cloze test and uses an LLM to fill in missing key terms to ensure precise key term matching, and (2) StepRestore, which shuffles ground truth steps and uses an LLM to reconstruct the actionable instructions in the correct order, verifying step order and completeness.To support comprehensive evaluation, we propose a benchmark dataset built upon the publicly available TechQA dataset, containing responses generated by different levels of QA systems. TechSupportEval achieves an AUC of 0.91, outperforming the state-of-the-art method by 7.6%. The code and dataset are available at https://github.com/NetManAIOps/TechSupportEval.
Yongqian Sun, Yuhe Liu, Longlong Xu, Zhe Xie, Changhua Pei, Fan Ni, Xuhui Cai, Dan Pei
IJCNN5
2025 Second-order Latent Factorization of Tensors based on Tucker Decomposition for spatio-temporal traffic flow data completion
abstract
The efficiency of Intelligent Transport Systems (ITS) runs on high-quality traffic data, however in real-world deployments, sensors failures, communication interruptions or other issues often lead to missing data, which affects the performance of ITS. Aiming at traffic data’s complex spatio-temporal characteristics, although the latent factorization of tensors (LFT) model has been widely used for missing-value completion, its non-convex objective function makes it difficult for first-order optimization methods to approximate high-quality second-order stationary points, therefore limiting the improvement of the completion accuracy. To address the issues, this paper proposes an incomplete tensor complementation model combining Tucker decomposition and second-order optimization strategy to improve the complementation accuracy and convergence stability. To address the issues, this paper proposes a Second-order Latent Factorization of Tensors based on Tucker Decomposition (SLTD), and efficiently solves it via Gauss-Newton approximation, so that it can significantly improve the model performance while keeping the computational cost low. Experimental results on real traffic datasets (in terms of average vehicle speed) from four cities verify the effectiveness of SLTD. Results show that the proposed model outperforms existing prevailing methods in terms of accuracy and provides a better solution for traffic data completion.
Jiajia Mi, Weiling Li, Huaqiang Yuan, Zhe Xie, Dongning Liu
SMC4
2025 ChatTS: Aligning Time Series with LLMs via Synthetic Data for Enhanced Understanding and Reasoning
abstract
Understanding time series is crucial for its application in real-world scenarios. Recently, large language models (LLMs) have been increasingly applied to time series tasks, leveraging their strong language capabilities to enhance various applications. However, research on multimodal LLMs (MLLMs) for time series understanding and reasoning remains limited, primarily due to the scarcity of high-quality datasets that align time series with textual information. This paper introduces ChatTS, a novel MLLM designed for time series analysis. ChatTS treats time series as a modality, similar to how vision MLLMs process images, enabling it to perform both understanding and reasoning with time series. To address the scarcity of training data, we propose an attribute-based method for generating synthetic time series and Time Series Evol-Instruct to generates diverse Q&As for enhanced reasoning capabilities. To the best of our knowledge, ChatTS is the first MLLM that takes multivariate time series as input for understanding and reasoning, which is fine-tuned exclusively on synthetic datasets. We evaluate its performance using benchmark datasets with real-world data, including six alignment tasks and four reasoning tasks. Our results show that ChatTS significantly outperforms existing vision-based MLLMs (e.g., GPT-4o) and text/agent-based LLMs, achieving a 46.0% improvement in alignment tasks and a 25.8% improvement in reasoning tasks. We have open-sourced the source code, model checkpoint and datasets at https://github.com/NetManAIOps/ChatTS.
Zhe Xie, Zeyan Li 0001, Xiao He 0008, Longlong Xu, Xidao Wen, Tieying Zhang, Jianjun Chen 0001, Dan Pei
Proc. VLDB Endow.1
2024 Microservice Root Cause Analysis With Limited Observability Through Intervention Recognition in the Latent Space
abstract
Many failure root cause analysis (RCA) algorithms for microservices have been proposed with the widespread adoption of microservices systems. Existing algorithms generally focus on RCA with ranking single-level (e.g. metric-level or service-level) root cause candidates (RCCs) with comprehensive monitoring metrics. However, many heterogeneous RCCs exist with limited observability in real-world microservices systems. Further, we find that the limited observability may result in inaccurate RCA through real-world failures in eBay. In this paper, for the first time, we propose to "model RCCs as latent variables". The core idea is to infer the status of RCCs as latent variables with related monitoring metrics instead of directly extracting features from only the observable metrics. Based on this, we propose LatentScope, an unsupervised RCA framework with heterogeneous RCCs under limited observability. A dual-space graph is proposed to model both observable and unobservable variables, with many-to-many relationships between spaces. To achieve fast inference of latent variables and RCA, we propose the LatentRegressor algorithm, which includes Regression-based Latent-space Intervention Recognition (RLIR) to achieve intervention recognition-based RCA in latent space. LatentScope has been deployed in eBay's production environment and evaluated on both eBay's real-world failures and a testbed dataset. The evaluation results show that, compared with baseline algorithms, our model significantly improves the Top-1 recall by 9.7%-57.9%. The source code of LatentScope and the dataset are available at https://github.com/NetManAIOps/LatentScope.
Zhe Xie, Shenglin Zhang, Yitong Geng, Yao Zhang 0009, Minghua Ma, Xiaohui Nie, Zhenhe Yao, Longlong Xu, Yongqian Sun, Dan Pei
KDD1
2023 Proximal Symmetric Non-negative Latent Factor Analysis: A Novel Approach to Highly-Accurate Representation of Undirected Weighted Networks
Yurong Zhong, Zhe Xie, Weiling Li, Xin Luo 0001
ICIC (4)2
2023 From Point-wise to Group-wise: A Fast and Accurate Microservice Trace Anomaly Detection Approach
abstract
As Internet applications continue to scale up, microservice architecture has become increasingly popular due to its flexibility and logical structure. Anomaly detection in traces that record inter-microservice invocations is essential for diagnosing system failures. Deep learning-based approaches allow for accurate modeling of structural features (i.e., call paths) and latency features (i.e., call response time), which can determine the anomaly of a particular trace sample. However, the point-wise manner employed by these methods results in substantial system detection overhead and impracticality, given the massive volume of traces (billion-level). Furthermore, the point-wise approach lacks high-level information, as identical sub-structures across multiple traces may be encoded differently. In this paper, we introduce the first Group-wise Trace anomaly detection algorithm, named GTrace. This method categorizes the traces into distinct groups based on their shared sub-structure, such as the entire tree or sub-tree structure. A group-wise Variational AutoEncoder (VAE) is then employed to obtain structural representations. Moreover, the innovative "predicting latency with structure" learning paradigm facilitates the association between the grouped structure and the latency distribution within each group. With the group-wise design, representation caching, and batched inference strategies can be implemented, which significantly reduces the burden of detection on the system. Our comprehensive evaluation reveals that GTrace outperforms state-of-the-art methods in both performances (2.64% to 195.45% improvement in AUC metrics and 2.31% to 40.92% improvement in best F-Score) and efficiency (21.9x to 28.2x speedup). We have deployed and assessed the proposed algorithm on eBay's microservices cluster, and our code is available at https://github.com/NetManAIOps/GTrace.git.
Zhe Xie, Changhua Pei, Wanxue Li, Huai Jiang, Liangfei Su, Gaogang Xie, Dan Pei
ESEC/SIGSOFT FSE1
2023 Multi-Constrained Symmetric Nonnegative Latent Factor Analysis for Accurately Representing Undirected Weighted Networks
abstract
An Undirected Weighted Network (UWN) is frequently encountered in a big-data-related application concerning the complex interactions among numerous nodes. A Symmetric High-Dimensional and Incomplete (SHDI) matrix can smoothly illustrate such a UWN, which contains rich knowledge like node interaction behaviors and local complexes. To extract desired knowledge from an SHDI matrix, an analysis model should carefully consider its topology for describing a UWN's intrinsic symmetry precisely. Representation learning to a UWN borrows the success of a pyramid of symmetry-aware models like a Symmetric Nonnegative Matrix Factorization (SNMF) model whose objective function utilizes a sole Latent Factor (LF) matrix for representing SHDI's symmetry precisely. However, they suffer from the following drawbacks: 1) their computational complexity is high; and 2) their modeling strategy narrows their representation features, making them suffer from low learning ability. Aiming at addressing the above critical issues, this paper proposes a Multi-constrained Symmetric Nonnegative Latent-factor-analysis (MSNL) model with two-fold ideas: 1) introducing multi-constraints composed of multiple LF matrices, i.e., inequality and equality ones into a data-density-oriented objective function for precisely representing the intrinsic symmetry of an SHDI matrix with broadened feature space; and 2) implementing an alternating direction method of multipliers (ADMM)-incorporated learning scheme for efficiently solving such a multi-constrained model. Empirical studies on three SHDI matrices from a real bioinformatics or industrial application demonstrate that the proposed MSNL model achieves higher representation accuracy than state-of-the-art models do, as well as promising computational efficiency.
Yurong Zhong, Zhe Xie, Weiling Li, Xin Luo 0001
SMC2
2023 Position-Aware Subgraph Neural Networks with Data-Efficient Learning
abstract
Data-efficient learning on graphs (GEL) is essential in real-world applications. Existing GEL methods focus on learning useful representations for nodes, edges, or entire graphs with "small" labeled data. But the problem of data-efficient learning for subgraph prediction has not been explored. The challenges of this problem lie in the following aspects: 1) It is crucial for subgraphs to learn positional features to acquire structural information in the base graph in which they exist. Although the existing subgraph neural network method is capable of learning disentangled position encodings, the overall computational complexity is very high. 2) Prevailing graph augmentation methods for GEL, including rule-based, sample-based, adaptive, and automated methods, are not suitable for augmenting subgraphs because a subgraph contains fewer nodes but richer information such as position, neighbor, and structure. Subgraph augmentation is more susceptible to undesirable perturbations. 3) Only a small number of nodes in the base graph are contained in subgraphs, which leads to a potential "bias" problem that the subgraph representation learning is dominated by these "hot" nodes. By contrast, the remaining nodes fail to be fully learned, which reduces the generalization ability of subgraph representation learning. In this paper, we aim to address the challenges above and propose a Position-Aware Data-Efficient Learning framework for subgraph neural networks called PADEL. Specifically, we propose a novel node position encoding method that is anchor-free, and design a new generative subgraph augmentation method based on a diffused variational subgraph autoencoder, and we propose exploratory and exploitable views for subgraph contrastive learning. Extensive experiment results on three real-world datasets show the superiority of our proposed method over state-of-the-art baselines.
Chang Liu 0078, Yuwen Yang, Zhe Xie, Hongtao Lu 0001, Yue Ding 0001
WSDM3
2023 Unsupervised Anomaly Detection on Microservice Traces through Graph VAE
abstract
The microservice architecture is widely employed in large Internet systems. For each user request, a few of the microservices are called, and a trace is formed to record the tree-like call dependencies among microservices and the time consumption at each call node. Traces are useful in diagnosing system failures, but their complex structures make it difficult to model their patterns and detect their anomalies. In this paper, we propose a novel dual-variable graph variational autoencoder (VAE) for unsupervised anomaly detection on microservice traces. To reconstruct the time consumption of nodes, we propose a novel dispatching layer. We find that the inversion of negative log-likelihood (NLL) appears for some anomalous samples, which makes the anomaly score infeasible for anomaly detection. To address this, we point out that the NLL can be decomposed into KL-divergence and data entropy, whereas lower-dimensional anomalies can introduce an entropy gap with normal inputs. We propose three techniques to mitigate this entropy gap for trace anomaly detection: Bernoulli & Categorical Scaling, Node Count Normalization, and Gaussian Std-Limit. On five trace datasets from a top Internet company, our proposed TraceVAE achieves excellent F-scores.
Zhe Xie, Wenxiao Chen, Wanxue Li, Huai Jiang, Liangfei Su, Dan Pei
WWW1
2022 Generic and Robust Performance Diagnosis via Causal Inference for OLTP Database Systems
abstract
Online transaction processing (OLTP) database systems provide an effective solution to data support for online applications with high concurrency and low latency. An interruption or performance degradation of OLTP database systems may impact the availability of services and bring substantial economic loss. Thus, diagnosing the issue timely and mitigating it rapidly are essential for database administrators (DBAs). However, performance diagnosis for database systems is challenging due to numerous abnormal metrics, complex failure propagation, and high-performance requirements. Existing works relying on anomaly detection or causal graph construction cannot handle all these challenges simultaneously. In this paper, we propose an unsupervised learning-based method, CauseRank, to perform root cause localization with superior efficiency, high accuracy, and good interpretability. Two key techniques in CauseRank are a novel causal discovery algorithm named Group-based Greedy Equivalent Search (G-GES) incorporated with domain knowledge which treats metric groups as nodes to capture failure propagation and a simple yet effective ranking method named Causal Oriented Personalized PageRank (COPP). Extensive experiments on 97 real-world failure cases collected from a large-scale Oracle database demonstrate the effectiveness of CauseRank, achieving 82.5% top-3 accuracy and 93.8% top-5 accuracy and outperforming baseline approaches. The core idea and framework of CauseRank are generic and can be applied to other large-scale system components.
Xianglin Lu, Zhe Xie, Zeyan Li 0001, Mingjie Li 0005, Xiaohui Nie, Nengwen Zhao, Qingyang Yu, Shenglin Zhang, Kaixin Sui, Dan Pei
CCGRID2
2022 A Dynamic Linear Bias Incorporation Scheme for Nonnegative Latent Factor Analysis
Yurong Zhong, Zhe Xie, Weiling Li, Xin Luo 0001
PRICAI (1)2
2021 Adversarial and Contrastive Variational Autoencoder for Sequential Recommendation
abstract
Sequential recommendation as an emerging topic has attracted increasing attention due to its important practical significance. Models based on deep learning and attention mechanism have achieved good performance in sequential recommendation. Recently, the generative models based on Variational Autoencoder (VAE) have shown the unique advantage in collaborative filtering. In particular, the sequential VAE model as a recurrent version of VAE can effectively capture temporal dependencies among items in user sequence and perform sequential recommendation. However, VAE-based models suffer from a common limitation that the representational ability of the obtained approximate posterior distribution is limited, resulting in lower quality of generated samples. This is especially true for generating sequences. To solve the above problem, in this work, we propose a novel method called Adversarial and Contrastive Variational Autoencoder (ACVAE) for sequential recommendation. Specifically, we first introduce the adversarial training for sequence generation under the Adversarial Variational Bayes (AVB) framework, which enables our model to generate high-quality latent variables. Then, we employ the contrastive loss. The latent variables will be able to learn more personalized and salient characteristics by minimizing the contrastive loss. Besides, when encoding the sequence, we apply a recurrent and convolutional structure to capture global and local relationships in the sequence. Finally, we conduct extensive experiments on four real-world datasets. The experimental results show that our proposed ACVAE model outperforms other state-of-the-art methods.
Zhe Xie, Chengxuan Liu, Hongtao Lu 0001, Dong Wang 0024, Yue Ding 0001
WWW1
2006 Web Services Based GIS Model Sharing Service
abstract
Nowadays, sharing is one of the most frequently hot topics. Compared with adequate data source and data share, it is more important to use the abundant data. Generally, a geographic information system (GIS) consists of two parts which are data and function. Sharing function is a good way to make the data usable. For this purpose, we set up the architecture for GIS model sharing. Furthermore, an archetype system is built to prove the feasibility, and an environment evaluation model is used for a test on Beijing and surrounding areas. The archetype system shares the environmental evaluation model via Web services. It enables end users to combine local data and remote function together to execute environment evaluation and analysis. The archetype system is a proof of the feasibility of the GIS model sharing architecture. The environment evaluation result generated by the archetype system is consistent with the ground truth. It is more convenient to use the archetype system and it also reveals the utility of Web services based GIS model sharing service. However, the archetype system needs to be improved in many aspects.
Zhongshi Tang, Zhe Xie, Hongrui Zhao
IGARSS2