VLDB 2026 Research / reviewers in the wild / expert
Min Li 0065
dblp:82/0-65
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0009-0006-0049-361XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LogGen: Integrating traditional model and LLM with code analysis for precise log generation
Min Li 0065, Gou Tan, Pengfei Chen 0002, Chuanfu Zhang |
J. Syst. Softw. | 1 |
| 2026 | Logfun: An efficient function-Level log management framework for systems implemented with python
Min Li 0065, Gou Tan, Mingdong He, Guangba Yu, Pengfei Chen 0002, Chuanfu Zhang |
J. Syst. Softw. | 1 |
| 2026 | LogBoost: Boost Log Anomaly Detection by Cherry-Picking Log SequencesabstractDebugging and operating services always benefit from logs. Since logs provide rich information on events and render comprehensive execution traces, it is imperative to automatically detect faults from extensive logs through log-based anomaly detection. However, due to the ineffectiveness and complexity of log-based detection models, they have not been widely adopted in template evaluation and real-world detection. Therefore, we propose LogBoost, a lightweight framework to boost log-based anomaly detection by automatically reducing redundant log templates. Based on our proposed similarity measurement, it effectively sorts the importance of log templates and identifies templates that are ineffective in anomaly detection. In evaluation, we introduce Spark-SDA, a new dataset featuring more diverse log templates and excessively long sequences, alongside the HDFS log dataset. We further evaluate LogBoost using established log-based anomaly detection models. The results demonstrate that eliminating ineffective log templates via LogBoost improves feature efficacy while reducing computational overhead. For instance, RandomForest obtains an F1-score at 0.983 given only 500 training samples on an optimized frequency vector of the HDFS dataset. Meanwhile, the lengths of log sequences are reduced by 51% to 77%, and the prediction time of deep learning models is reduced by 55% to 81%. Our results show that LogBoost is an effective approach to accelerate log-based anomaly detection. Min Li 0065, Pengfei Chen 0002, Yuanhao Lai, Zibin Zheng |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | FaaSScout: Fast and Full Lifecycle RCA for FaaS Applications Using Salient Feature Mining
Min Li 0065, Pengfei Chen 0002, Chongkang Tan |
IEEE Trans. Cloud Comput. | 1 |
| 2024 | A Bayesian LSTM Based Active Anomaly Detection Service for Large Online SystemsabstractCurrently, many large online systems are constructed with a microservice architecture. Due to the complex dependencies, the failure of a service in such a system can cause an avalanche, which directly affects user experience and the company’s revenue. It is critical for service operators to build anomaly detection services to monitor online systems closely and comprehensively. Even though a large number of anomaly detection approaches have been proposed, few of them can simultaneously adapt to hundreds of operators’ practical detection requirements. To tackle this problem, we proposed LSTM-AAD, a Bayesian LSTM based active anomaly detection service. LSTM-AAD extracts anomaly features based on the common patterns among metrics, introduces a Bayesian LSTM model to detect anomalies in time series metrics, and employs active learning to update the online model via a small number of uncertain feedback samples. In addition, the proposed user-oriented service can be quickly responsive to operators’ further requirements. We conduct extensive experiments on real time series metrics of large online services in Tencent. The results indicate that LSTM-AAD significantly outperforms other state-of-the-art methods. Moreover, our approach can detect anomalies efficiently out of box to work in a large-scale system. Chen Wang 0075, Tao Huang 0021, Min Li 0065, Pengfei Chen 0002 |
Internetware | 3 |
| 2023 | Online Data Drift Detection for Anomaly Detection Services based on Deep Learning towards Multivariate Time SeriesabstractDeep learning models have been successfully adopted in anomaly detection for multivariate time series data in various fields. These models are good at capturing complex time dependencies and extracting meaningful patterns from time series data. However, the trained models may become outdated due to unforeseen changes in real-world data, which can lead to a decrease in the quality of model service. Therefore, it is crucial to continuously monitor the performance of the model and analyze its behavior to ensure its reliability and availability. We propose an online data drift detection method that uses an unsupervised deep learning network, Variational Autoencoder (VAE), to monitor deep learning models in the field of multivariate time series anomaly detection. This method consists of three main steps namely data collection and statistical analysis, real-time drift detection, and drift interpretation. We collect raw time series data and model prediction data non-invasively from the model server. Then they are separated into windows for drift detection. Furthermore, the method can provide analysis and interpretation when drift is detected. Our evaluation experiments involve three real-world datasets from various industrial domains and four different structured anomaly detection models. We validate the effectiveness of drift detection in multivariate time series, and then test how the anomaly detection models perform during data drift detection. The highest improvement in F1 score is approximately 0.16. In addition, we provide an analysis of the interpretability of the model performance. Gou Tan, Pengfei Chen 0002, Min Li 0065 |
QRS | 3 |