Haotian Si

dblp:355/0412 · DBLP profile ↗
← Back
6ranked-venue papers
3as first author
6since 2021 · last 2026
0009-0006-8862-7781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Software engineering, systems software and programming languages · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Smart Eye: LLM-Guided Proposer-Verifier Framework for Industrial-Scale Log Anomaly Detection
Changhua Pei, Hang Cui 0004, Xinyuan Liao, Cenjie Hu, Haotian Si, Ke Xiang, Gaogang Xie, Dan Pei
WWW11
2026 ViTs: Teaching Machines to See Time Series Anomalies Like Human Experts
abstract
Web service administrators must ensure the stability of multiple systems by promptly detecting anomalies in Key Performance Indicators (KPIs). Achieving the goal of "train once, infer across scenarios" remains a fundamental challenge for time series anomaly detection models. Beyond improving zero-shot generalization, such models must also flexibly handle sequences of varying lengths during inference, ranging from one hour to one week, without retraining. Conventional approaches rely on sliding-window encoding and self-supervised learning, which restrict inference to fixed-length inputs. Large Language Models (LLMs) have demonstrated remarkable zero-shot capabilities across general domains. However, when applied to time series data, they face inherent limitations due to context length. To address this issue, we propose ViTs, a Vision-Language Model (VLM)-based framework that converts time series curves into visual representations. By rescaling time series images, temporal dependencies are preserved while maintaining a consistent input size, thereby enabling efficient processing of arbitrarily long sequences without context constraints. Training VLMs for this purpose introduces unique challenges, primarily due to the scarcity of aligned time series image-text data. To overcome this, we employ an evolutionary algorithm to automatically generate thousands of high-quality image-text pairs and design a three-stage training pipeline consisting of: (1) time series knowledge injection, (2) anomaly detection enhancement, and (3) anomaly reasoning refinement. Extensive experiments demonstrate that ViTs substantially enhance the ability of VLMs to understand and detect anomalies in time series data. All datasets and code will be publicly released at: https://anonymous.4open.science/r/ViTs-C484/.
Changhua Pei, Yang Liu 0442, Hengyue Jiang, Haotian Si, Hang Cui 0004, Gaogang Xie, Dan Pei
WWW6
2025 CMoS: Rethinking Time Series Prediction Through the Lens of Chunk-wise Spatial Correlations
abstract
Recent advances in lightweight time series forecasting models suggest the inherent simplicity of time series forecasting tasks. In this paper, we present CMoS, a super-lightweight time series forecasting model. Instead of learning the embedding of the shapes, CMoS directly models the spatial correlations between different time series chunks. Additionally, we introduce a Correlation Mixing technique that enables the model to capture diverse spatial correlations with minimal parameters, and an optional Periodicity Injection technique to ensure faster convergence. Despite utilizing as low as 1% of the lightweight model DLinear’s parameters count, experimental results demonstrate that CMoS outperforms existing state-of-the-art models across multiple datasets. Furthermore, the learned weights of CMoS exhibit great interpretability, providing practitioners with valuable insights into temporal structures within specific application scenarios.
Haotian Si, Changhua Pei, Dan Pei, Gaogang Xie
ICML1
2025 DeST: An Unsupervised Decoupled Spatio-Temporal Framework for Microservice Incident Management
abstract
Effective incident management in large-scale microservice systems demands both accurate anomaly detection (AD) and precise root cause localization (RCL) across heterogeneous data modalities. However, existing approaches often treat these tasks in isolation, resulting in redundant maintenance, delayed response, and the absence of shared diagnostic context. While recent efforts have explored unified frameworks to support both tasks, these approaches often suffer from high falsealarm rates due to cross-modal interference. To address these issues, we propose DeST, an unsupervised decoupled spatiotemporal framework that jointly performs anomaly detection and root cause localization. DeST proposes a multi-stage fusion strategy that decouples temporal and spatial feature learning to mitigate cross-modal interference and prevent cross-modal interference. Furthermore, it incorporates task-specific modal routing to direct learned representations to different tasks, enhancing both detection and localization accuracy. To ensure robustness against transient noise, DeST designs a Differential Multi-Scale Convolutional Network (DMCN) for noise-resistant temporal feature representation. We evaluate DeST on two real-world microservice benchmarks, where it achieves a perfect F1-score of $\mathbf{1. 0 0}$ for anomaly detection and outperforms existing methods in root cause localization accuracy. Ablation studies highlight the effectiveness of key components. Our unified framework reduces false alarms in anomaly detection and streamlines root cause localization, providing a robust and practical solution for microservice incident management.
Xiaohui Nie, Hang Cui 0004, Changhua Pei, Haotian Si, Ke Xiang, Yanbiao Li 0001, Gaogang Xie, Dan Pei
ISSRE4
2024 TimeSeriesBench: An Industrial-Grade Benchmark for Time Series Anomaly Detection Models
abstract
Time series anomaly detection (TSAD) has gained significant attention due to its real-world applications to improve the stability of modern software systems. However, there is no effective way to verify whether they can meet the requirements for real-world deployment. Firstly, current algorithms typically train a specific model for each time series. Maintaining such many models is impractical in a large-scale system with tens of thousands of curves. The performance of using merely one unified model to detect anomalies remains unknown. Secondly, most TSAD models are trained on the historical part of a time series and are tested on its future segment. In distributed systems, however, there are frequent system deployments and upgrades, with new, previously unseen time series emerging daily. The performance of testing newly incoming unseen time series on current TSAD algorithms remains unknown. Lastly, the assumptions of the evaluation metrics in existing benchmarks are far from practical demands. To solve the above-mentioned problems, we propose an industrial-grade benchmark TimeSeriesBench. We assess the performance of existing algorithms across more than 168 evaluation settings and provide comprehensive analysis for the future design of anomaly detection algorithms. An industrial dataset is also released along with TimeSeriesBench.
Haotian Si, Changhua Pei, Hang Cui 0004, Yongqian Sun, Shenglin Zhang, Haiming Zhang 0002, Dan Pei, Gaogang Xie
ISSRE1
2023 Beyond Sharing: Conflict-Aware Multivariate Time Series Anomaly Detection
abstract
Massive key performance indicators (KPIs) are monitored as multivariate time series data (MTS) to ensure the reliability of the software applications and service system. Accurately detecting the abnormality of MTS is very critical for subsequent fault elimination. The scarcity of anomalies and manual labeling has led to the development of various self-supervised MTS anomaly detection (AD) methods, which optimize an overall objective/loss encompassing all metrics' regression objectives/losses. However, our empirical study uncovers the prevalence of conflicts among metrics' regression objectives, causing MTS models to grapple with different losses. This critical aspect significantly impacts detection performance but has been overlooked in existing approaches. To address this problem, by mimicking the design of multi-gate mixture-of-experts (MMoE), we introduce CAD, a Conflict-aware multivariate KPI Anomaly Detection algorithm. CAD offers an exclusive structure for each metric to mitigate potential conflicts while fostering inter-metric promotions. Upon thorough investigation, we find that the poor performance of vanilla MMoE mainly comes from the input-output misalignment settings of MTS formulation and convergence issues arising from expansive tasks. To address these challenges, we propose a straightforward yet effective task-oriented metric selection and p&s (personalized and shared) gating mechanism, which establishes CAD as the first practicable multi-task learning (MTL) based MTS AD model. Evaluations on multiple public datasets reveal that CAD obtains an average F1-score of 0.943 across three public datasets, notably outperforming state-of-the-art methods. Our code is accessible at https://github.com/dawnvince/MTS_CAD.
Haotian Si, Changhua Pei, Zhihan Li 0002, Haiming Zhang 0002, Zulong Diao, Gaogang Xie, Dan Pei
ESEC/SIGSOFT FSE1