VLDB 2026 Research / reviewers in the wild / expert
Zhenyu Zhong
dblp:36/5105
· DBLP profile ↗
21ranked-venue papers
4as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Software engineering, systems software and programming languages · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Systems, architecture and hardware · 3 · 2 first-authorComputer networks · 2 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Revisiting the Reliability of Language Models in Instruction-FollowingabstractAdvanced LLMs have achieved near-ceiling instruction-following accuracy on benchmarks such as IFEVAL.However, these impressive scores do not necessarily translate to reliable services in real-world use, where users often vary their phrasing, contextual framing, and task formulations.In this paper, we study nuance-oriented reliability: whether models exhibit consistent competence across cousin prompts that convey analogous user intents but with subtle nuances.To quantify this, we introduce a new metric, reliable@k, and develop an automated pipeline that generates high-quality cousin prompts via data augmentation.Building upon this, we construct IFE-VAL++ for systematic evaluation.Across 20 proprietary and 26 open-source LLMs, we find that current models exhibit substantial insufficiency in nuance-oriented reliability-their performance can drop by up to 61.8% with nuanced prompt modifications.What's more, we characterize it and explore three potential improvement recipes.Our findings highlight nuance-oriented reliability as a crucial yet underexplored next step toward more dependable and trustworthy LLM behavior.Our code and benchmark are accessible: https: //github.com/jianshuod/IFEval-pp. Jianshuo Dong, Liu Yan, Zhenyu Zhong, Tao Wei 0002, Chao Zhang 0008, Han Qiu 0001 |
ACL (1) | 4 |
| 2026 | Failure Diagnosis in Microservice Systems: A Comprehensive Survey and AnalysisabstractWidely adopted for their scalability and flexibility, modern microservice systems present unique failure diagnosis challenges due to their independent deployment and dynamic interactions. This complexity can lead to cascading failures that negatively impact operational efficiency and user experience. Recognizing the critical role of fault diagnosis in improving the stability and reliability of microservice systems, researchers have conducted extensive studies and achieved a number of significant results. This survey provides an exhaustive review of 98 scientific papers from 2003 to the present, including a thorough examination and elucidation of the fundamental concepts, system architecture, and problem statement. It also includes a qualitative analysis of the dimensions, providing an in-depth discussion of current best practices and future directions, aiming to further its development and application. In addition, this survey compiles publicly available datasets, toolkits, and evaluation metrics to facilitate the selection and validation of techniques for practitioners. Shenglin Zhang, Sibo Xia, Wenzhao Fan, Binpeng Shi, Xiao Xiong, Zhenyu Zhong, Minghua Ma, Yongqian Sun, Dan Pei |
ACM Trans. Softw. Eng. Methodol. | 6 |
| 2026 | LLM-Enhanced Failure Localization in Microservices: Integrating Multi-Modal Data and Expert InterpretationabstractFailure localization in microservice environments is increasingly challenging. While large language models (LLMs) have shown promise in software engineering tasks, existing approaches struggle to effectively integrate multi-modal telemetry data (e.g., log, metric and trace) and provide interpretable results. This paper presents LocaleXpert, a novel failure localization system that combines specialized LLM-based agents with traditional AIOps methods to diagnose issues in microservice environments. LocaleXpert introduces three key innovations: (1) a modular pipeline that transforms metrics, logs, and traces into natural language descriptions that LLMs can effectively process, (2) specialized expert agents that analyze each data type and collaborate to identify root causes, and (3) an interpretation mechanism that produces clear, actionable explanations of its reasoning process. Evaluation results show that LocaleXpert significantly outperforming baseline approaches in both accuracy and interpretability. The system has been successfully deployed in Microsoft's AIOpsLab benchmark, demonstrating its effectiveness. Zhenyu Zhong, Ruowei Fu, Minghua Ma, Shenglin Zhang, Yongqian Sun, Chetan Bansal, Dan Pei |
IEEE Trans. Serv. Comput. | 1 |
| 2025 | "I've Decided to Leak": Probing Internals Behind Prompt Leakage IntentsabstractLarge language models (LLMs) exhibit prompt leakage vulnerabilities, where they may be coaxed into revealing system prompts embedded in LLM services, raising intellectual property and confidentiality concerns.An intriguing question arises: Do LLMs genuinely internalize prompt leakage intents in their hidden states before generating tokens?In this work, we use probing techniques to capture LLMs' intent-related internal representations and confirm that the answer is yes.We start by comprehensively inducing prompt leakage behaviors across diverse system prompts, attack queries, and decoding methods.We develop a hybrid labeling pipeline, enabling the identification of broader prompt leakage behaviors beyond mere verbatim leaks.Our results show that a simple linear probe can predict prompt leakage risks from pre-generation hidden states without generating any tokens.Across all tested models, linear probes consistently achieve 90%+ AUROC, even when applied to new system prompts and attacks.Understanding the model internals behind prompt leakage drives practical applications, including intention-based detection of prompt leakage risks. Jianshuo Dong, Liu Yan, Zhenyu Zhong, Tao Wei 0002, Ke Xu 0002, Minlie Huang, Chao Zhang 0008, Han Qiu 0001 |
EMNLP | 4 |
| 2025 | Harnessing Massive Satellite Imagery with Efficient Masked Image Modeling
Fengxiang Wang 0004, Hongzhen Wang, Di Wang 0023, Zonghao Guo, Zhenyu Zhong, Long Lan, Wenjing Yang 0002, Jing Zhang 0037 |
ICCV | 5 |
| 2025 | LLM-Powered Multi-Agent Collaboration for Intelligent Industrial On-Call AutomationabstractIn large-scale enterprises, on-call engineers (OCEs) are critical for ensuring service availability and reliability. However, as incidents grow in volume and complexity, traditional manual on-call processes are becoming increasingly inadequate. Recent advances in large language models (LLMs) have demonstrated remarkable capabilities in reasoning and multi-agent collaboration, presenting new opportunities for automation. We propose OncallX, an end-to-end automated on-call system designed for real-world industrial scenarios that integrates LLMs with multi-agent cooperation to enable intelligent and efficient incident management. OncallX first enhances user queries by leveraging external knowledge bases and multi-turn dialogue interactions. Subsequently, multiple expert agents collaborate through tree-search-based mechanisms to generate effective responses and solutions. When incidents cannot be resolved automatically, OncallX accurately assigns them to the most appropriate teams. Comprehensive experiments conducted in the real-world production environment of a top-tier global online video service provider demonstrate that OncallX efficiently responds to incidents and accurately triages tickets, significantly outperforming existing methods in both automated metrics and human evaluations. Furthermore, OncallX has been successfully deployed in production for two months, during which it has substantially enhanced on-call efficiency, reducing average incident response time to just 21 seconds and average triage time to 4 seconds—representing a transformative improvement in operational excellence. Ruowei Fu, Yang Zhang 0103, Zeyu Che, Zhenyu Zhong, Zhiqiang Ren, Shenglin Zhang, Feng Wang 0054, Yongqian Sun, Yu Zhang 0209 |
ASE | 5 |
| 2025 | Interpretable Failure Localization for Microservice Systems Based on Graph AutoencoderabstractAccurate and efficient localization of root cause instances in large-scale microservice systems is of paramount importance. Unfortunately, prevailing methods face several limitations. Notably, some recent methods rely on supervised learning which necessitates a substantial amount of labeled data. However, labeling root cause instances is time-consuming and laborious, especially with multiple modalities of data including logs, traces, metrics, and so on. Moreover, some approaches favor deep learning for localization but lack interpretability and continuous improvement mechanisms. To address the above challenges, we propose DeepHunt , a novel root cause localization method based on multimodal data analysis. Firstly, DeepHunt introduces root cause score (RCS) by integrating reconstruction errors and failure propagation patterns (upstream–downstream relationships), imparting interpretability to the localization of root causes. Then, it embraces graph autoencoder (GAE) to address the limitation imposed by scarce labeled data. It employs data augmentation to mitigate the adverse effects of insufficient historical training samples. We evaluate DeepHunt on two open source datasets, and it outperforms existing methods when facing a zero-label cold start. DeepHunt can be further improved by continuously fine-tuning through a feedback mechanism. Yongqian Sun, Binpeng Shi, Shenglin Zhang, Shiyu Ma, Pengxiang Jin, Zhenyu Zhong, Lemeng Pan, Yicheng Guo, Dan Pei |
ACM Trans. Softw. Eng. Methodol. | 7 |
| 2023 | An Empirical Analysis of Anomaly Detection Methods for Multivariate Time SeriesabstractUsing multivariate time series (MTS) data for anomaly detection is widely adopted in service systems, such as web services and financial businesses. Researchers have recently proposed some well-performed algorithms for MTS anomaly detection from different perspectives. When applied to the real world, we observe that none of the algorithms is adaptable to all scenarios due to the complex data and anomaly characteristics. Moreover, there is currently a lack of comprehensive analysis work of these algorithms to guide operators in selecting the appropriate one in practice. To bridge this gap, we conduct an empirical study using various real-world data to gain an in-depth understanding of state-of-the-art anomaly detection algorithms. First, we provide general recommendations to guide operators in selecting suitable models based on the volume of training data, computational resources, and effectiveness requirements. Then, we summarize the typical data characteristics and types of anomalies and offer tailored model selection suggestions for different data characteristics and anomaly types. At last, we apply the summarized model selection suggestions to all the datasets we collected. The results show that most of our suggestions can achieve better than any single algorithm alone, demonstrating the effectiveness and generalization of our recommendations. Dongwen Li, Shenglin Zhang, Yongqian Sun, Zeyu Che, Zhenyu Zhong, Minghan Liang, Minyi Shao, Mingjie Li 0005, Dan Pei |
ISSRE | 7 |
| 2023 | Robust Multimodal Failure Detection for Microservice SystemsabstractProactive failure detection of instances is vitally essential to microservice systems because an instance failure can propagate to the whole system and degrade the system's performance. Over the years, many single-modal (i.e., metrics, logs, or traces) databased anomaly detection methods have been proposed. However, they tend to miss a large number of failures and generate numerous false alarms because they ignore the correlation of multimodal data. In this work, we propose AnoFusion, an unsupervised failure detection approach, to proactively detect instance failures through multimodal data for microservice systems. It applies a Graph Transformer Network (GTN) to learn the correlation of the heterogeneous multimodal data and integrates a Graph Attention Network (GAT) with Gated Recurrent Unit (GRU) to address the challenges introduced by dynamically changing multimodal data. We evaluate the performance of AnoFusion through two datasets, demonstrating that it achieves the F1-score of 0.857 and 0.922, respectively, outperforming the state-of-the-art failure detection approaches. Minghua Ma, Zhenyu Zhong, Shenglin Zhang, Zhiyuan Tan 0005, Xiao Xiong, LuLu Yu, Yongqian Sun, Dan Pei, Qingwei Lin, Dongmei Zhang 0001 |
KDD | 3 |
| 2023 | Robust Failure Diagnosis of Microservice System Through Multimodal DataabstractAutomatic failure diagnosis is crucial for large microservice systems. Currently, most failure diagnosis methods rely solely on single-modal data (i.e., using either metrics, logs, or traces). In this study, we conduct an empirical study using real-world failure cases to show that combining these sources of data (multimodal data) leads to a more accurate diagnosis. However, effectively representing these data and addressing imbalanced failures remain challenging. To tackle these issues, we proposeDiagFusion, a robust failure diagnosis approach that uses multimodal data. It leverages embedding techniques and data augmentation to represent the multimodal data of service instances, combines deployment data and traces to build a dependency graph, and uses a graph neural network to localize the root cause instance and determine the failure type. Our evaluations using real-world datasets show thatDiagFusionoutperforms existing methods in terms of root cause instance localization (improving by 20.9% to 368%) and failure type determination (improving by 11.0% to 169%). Shenglin Zhang, Pengxiang Jin, Yongqian Sun, Bicheng Zhang, Sibo Xia, Zhengdan Li, Zhenyu Zhong, Minghua Ma, Wa Jin, Dan Pei |
IEEE Trans. Serv. Comput. | 8 |
| 2022 | Detecting multi-sensor fusion errors in advanced driver-assistance systemsabstractAdvanced Driver-Assistance Systems (ADAS) have been thriving and widely deployed in recent years. In general, these systems receive sensor data, compute driving decisions, and output control signals to the vehicles. To smooth out the uncertainties brought by sensor outputs, they usually leverage multi-sensor fusion (MSF) to fuse the sensor outputs and produce a more reliable understanding of the surroundings. However, MSF cannot completely eliminate the uncertainties since it lacks the knowledge about which sensor provides the most accurate data and how to optimally integrate the data provided by the sensors. As a result, critical consequences might happen unexpectedly. In this work, we observed that the popular MSF methods in an industry-grade ADAS can mislead the car control and result in serious safety hazards. We define the failures (e.g., car crashes) caused by the faulty MSF as fusion errors and develop a novel evolutionary-based domain-specific search framework, FusED, for the efficient detection of fusion errors. We further apply causality analysis to show that the found fusion errors are indeed caused by the MSF method. We evaluate our framework on two widely used MSF methods in two driving environments. Experimental results show that FusED identifies more than 150 fusion errors. Finally, we provide several suggestions to improve the MSF methods we study. Ziyuan Zhong, Zhisheng Hu, Shengjian Guo, Xinyang Zhang 0001, Zhenyu Zhong, Baishakhi Ray |
ISSTA | 5 |
| 2022 | Robust System Instance Clustering for Large-Scale Web ServicesabstractSystem instance clustering is crucial for large-scale Web services because it can significantly reduce the training overhead of anomaly detection methods. However, the vast number of system instances with massive time points, redundant metrics, and noise bring significant challenges. We propose OmniCluster to accurately and efficiently cluster system instances for large-scale Web services. It combines a one-dimensional convolutional autoencoder (1D-CAE), which extracts the main features of system instances, with a simple, novel, yet effective three-step feature selection strategy. We evaluated OmniCluster using real-world data collected from a top-tier content service provider providing services for one billion+ monthly active users (MAU), proving that OmniCluster achieves high accuracy (NMI=0.9160) and reduces the training overhead of five anomaly detection models by 95.01% on average. Shenglin Zhang, Dongwen Li, Zhenyu Zhong, Minghan Liang, Jiexi Luo, Yongqian Sun, Ya Su, Sibo Xia, Zhongyou Hu, Dan Pei, Jiyan Sun, Yinlong Liu |
WWW | 3 |
| 2022 | Efficient KPI Anomaly Detection Through Transfer Learning for Large-Scale Web ServicesabstractTimely anomaly detection of key performance indicators (KPIs),e.g., service response time, error rate, is of utmost importance to Web services. Over the years, many unsupervised deep learning-based anomaly detection approaches have been proposed. To achieve good performance, they require a long period of KPI data for model training, which is not easy to guarantee with frequent service changes. Additionally, the training overhead is too significant for the vast number of KPIs in large-scale Web services. To address the problems, we propose an unsupervised KPI anomaly detection approach, namedAnoTransfer, by combining a novel Variational Auto-Encoder (VAE)-based KPI clustering algorithm with an adaptive transfer learning strategy. Extensive evaluation experiments using real-world data collected from several large-scale Web service providers demonstrate thatAnoTransferreduces the average initialization time by 65.71% and improves the training efficiency by 50.62 times, without significantly degrading anomaly detection accuracy. Shenglin Zhang, Zhenyu Zhong, Dongwen Li, Qiliang Fan, Yongqian Sun, Man Zhu, Dan Pei, Jiyan Sun, Yinlong Liu, Yongqiang Zou |
IEEE J. Sel. Areas Commun. | 2 |
| 2022 | Potential escalator-related injury identification and prevention based on multi-module integrated system for public health
Zeyu Jiao, Huan Lei, Hengshan Zong, Yingjie Cai, Zhenyu Zhong |
Mach. Vis. Appl. | 5 |
| 2022 | Longitudinal Platoon Control of Connected Vehicles: Analysis and VerificationabstractThis paper proposes a longitudinal platoon controller for connected vehicles (CVs) by considering the information of multiple preceding vehicles and the car-following interactions between CVs. The stability of the proposed controller is analyzed using the Routh criterion. For the verification, we develop an integrated platoon control framework for CVs in a V2V/V2I communication environment. The proposed framework consists of two main components: simulation platform and experimental platform. In particular, the simulation platform is developed based on the TransModeler software, and the experimental platform is designed using the self-developed V2X devices. Finally, a scenario of platoon forming is taken as an example and is conducted in simulation platform and experimental platform, respectively. Results demonstrate the effectiveness of the proposed controller with respect to the trajectory and velocity profiles. Yongfu Li 0001, Zhenyu Zhong, Qi Sun 0004, Simon Hu 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2020 | Quantifying DNN Model Robustness to the Real-World ThreatsabstractDNN models have suffered from adversarial example attacks, which lead to inconsistent prediction results. As opposed to the gradient-based attack, which assumes white-box access to the model by the attacker, we focus on more realistic input perturbations from the real-world and their actual impact on the model robustness without any presence of the attackers. In this work, we promote a standardized framework to quantify the robustness against real-world threats. It is composed of a set of safety properties associated with common violations, a group of metrics to measure the minimal perturbation that causes the offense, and various criteria that reflect different aspects of the model robustness. By revealing comparison results through this framework among 13 pre-trained ImageNet classifiers, three state-of-the-art object detectors, and three cloud-based content moderators, we deliver the status quo of the real-world model robustness. Beyond that, we provide robustness benchmarking datasets for the community. Zhenyu Zhong, Zhisheng Hu |
DSN | 1 |
| 2020 | Fooling Detection Alone is Not Enough: Adversarial Attack against Multiple Object Tracking
Yunhan Jia, Yantao Lu, Junjie Shen 0001, Qi Alfred Chen, Hao Chan, Zhenyu Zhong, Tao Wei 0002 |
ICLR | 6 |
| 2011 | Speed Up Statistical Spam Filter by ApproximationabstractStatistical-based Bayesian filters have become a popular and important defense against spam. However, despite their effectiveness, their greater processing overhead can prevent them from scaling well for enterprise level mail servers. For example, the dictionary lookups that are characteristic of this approach are limited by the memory access rate, therefore relatively insensitive to increases in CPU speed. We conduct a comprehensive study to address this scaling issue by proposing a series of acceleration techniques that speed up Bayesian filters based on approximate classifications. The core approximation technique uses hash-based lookup and lossy encoding. Lookup approximation is based on the popular Bloom filter data structure with an extension to support value retrieval. Lossy encoding is used to further compress the data structure. While these approximation methods introduce additional errors to a strict Bayesian approach, we show how the errors can be both minimized and biased toward a false negative classification. We demonstrate a 6× speedup over two well-known spam filters (bogofilter and qsf) while achieving an identical false positive rate and similar false negative rate to the original filters. Zhenyu Zhong, Kang Li 0001 |
IEEE Trans. Computers | 1 |
| 2010 | Mining DNS for malicious domain registrationsabstractMillions of new domains are registered every day and the many of them are malicious. It is challenging to keep track of malicious domains by only Web content analysis due to the large number of domains. One interesting pattern in legitimate domain names is that many of them consist of English words Yuanchen He, Zhenyu Zhong, Sven Krasser, Yuchun Tang |
CollaborateCom | 2 |
| 2009 | Privacy-Aware Collaborative Spam FilteringabstractWhile the concept of collaboration provides a natural defense against massive spam e-mails directed at large numbers of recipients, designing effective collaborative anti-spam systems raises several important research challenges. First and foremost, since e-mails may contain confidential information, any collaborative anti-spam approach has to guarantee strong privacy protection to the participating entities. Second, the continuously evolving nature of spam demands the collaborative techniques to be resilient to various kinds of camouflage attacks. Third, the collaboration has to be lightweight, efficient, and scalable. Toward addressing these challenges, this paper presents ALPACAS-a privacy-aware framework for collaborative spam filtering. In designing the ALPACAS framework, we make two unique contributions. The first is a feature-preserving message transformation technique that is highly resilient against the latest kinds of spam attacks. The second is a privacy-preserving protocol that provides enhanced privacy guarantees to the participating entities. Our experimental results conducted on a real e-mail data set shows that the proposed framework provides a 10 fold improvement in the false negative rate over the Bayesian-based Bogofilter when faced with one of the recent kinds of spam attacks. Further, the privacy breaches are extremely rare. This demonstrates the strong privacy protection provided by the ALPACAS system. Kang Li 0001, Zhenyu Zhong, Lakshmish Ramaswamy |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2008 | ALPACAS: A Large-Scale Privacy-Aware Collaborative Anti-Spam SystemabstractWhile the concept of collaboration provides a natural defense against massive spam emails directed at large numbers of recipients, designing effective collaborative anti-spam systems raises several important research challenges. First and foremost, since emails may contain confidential information, any collaborative anti-spam approach has to guarantee strong privacy protection to the participating entities. Second, the continuously evolving nature of spam demands the collaborative techniques to be resilient to various kinds of camouflage attacks. Third, the collaboration has to be lightweight, efficient, and scalable. Towards addressing these challenges, this paper presents ALPACAS - a privacy-aware framework for collaborative spam filtering. In designing the ALPACAS framework, we make two unique contributions. The first is a feature-preserving message transformation technique that is highly resilient against the latest kinds of spam attacks. The second is a privacy-preserving protocol that provides enhanced privacy guarantees to the participating entities. Our experimental results conducted on a real email dataset shows that the proposed framework provides a 10 fold improvement in the false negative rate over the Bayesian-based Bogofilter when faced with one of the recent kinds of spam attacks. Further, the privacy breaches are extremely rare. This demonstrates the strong privacy protection provided by the ALPACAS system. Zhenyu Zhong, Lakshmish Ramaswamy, Kang Li 0001 |
INFOCOM | 1 |