VLDB 2026 Research / reviewers in the wild / expert
Haocheng Zhong
dblp:257/2947
· DBLP profile ↗
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0005-4975-668XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PaTGen: Temporal Similarity-Driven Proxy Benchmark Generation Method for Cloud WorkloadsabstractThe rapid expansion of cloud computing has made precise performance evaluation a critical necessity. However, conventional cloud benchmarks often face significant limitations in simulation environments—necessary for scalable and cost-effective testing—due to the complexity of technology stacks and substantial runtime overheads. Proxy benchmarking has thus emerged as a practical alternative. Existing methods primarily focus on the global similarity of micro-architectural metrics between proxy benchmarks and real workloads but neglect their temporal similarity, leading to inaccurate performance evaluations, flawed cache behavior simulations, and misguided architectural optimization decisions. To address this, we present PaTGen , a phase-aware method for generating proxy benchmarks that accurately reflect both global and temporal similarity. By partitioning workloads into phases and formulating proxy generation as nonlinear optimization problems, PaTGen further refines intra-phase execution patterns via the delay-based temporal similarity optimization (DTSO) technique. Evaluations on 15 real-world workloads show PaTGen achieves over 97% global similarity in key metrics while significantly outperforming state-of-the-art methods in temporal similarity. Ablation studies confirm the efficacy of phase division and DTSO. Further experiments confirm its scalability and generalizability across architectures. Moreover, the effectiveness observed in downstream tasks provides empirical evidence that preserving temporal similarity is a fundamental requirement for proxy benchmarks to faithfully capture real workload behavior. Haolang Yin, Weiwei Lin 0001, Huikang Huang, Xiaoxuan Luo, Haocheng Zhong, Keqin Li 0001 |
ACM Trans. Archit. Code Optim. | 5 |
| 2025 | AMORA: An Advanced Malleable and Operational Framework for Performance Prediction of Big Data SystemsabstractABSTRACT Background In the data era, big data systems have emerged as pivotal tools, underscoring the importance of performance prediction in enhancing the efficiency of big data clusters. Numerous performance models have been proposed, often grounded in artificial intelligence or simulation methodologies. While the bulk of research focuses on refining prediction precision and minimizing overhead, limited attention has been given to the consignation and standardization of these models. Objectives To bridge this gap between model developers and end‐users, this paper introduces AMORA—a novel versatile framework tailored for predicting the performance of big data systems. Methods Leveraging the identified behavior descriptions‐computation submodels (BD‐CS) pattern that is prevalent among various big data job performance models, AMORA allows access to different plugins accommodating different performance models' implementations. This framework also integrates a novel mutable computation graph technique to facilitate backtracking computation. Furthermore, AMORA's functionality extends to comprehensive end‐to‐end usability by enabling the acceptance of origin configuration files from diverse big data systems and presenting easily interpretable prediction reports. Results This work demonstrates AMORA's efficacy in producing an accurate trace of Hadoop job through the selection of appropriate performance model plugins and parameter adjustments and showcasing the application of the proposed mutable computation graph technique in calculating the starting moment of an early‐start reducer. Additionally, two validation experiments are conducted, involving the implementation of various Hadoop and Spark performance models, respectively. The experiment results manifest the prediction precision and overheads of these performance models. Conclusion These experiments exhibit AMORA's role as a benchmark platform for implementing various types of big data job performance models catered to diverse big data systems. Weiwei Lin 0001, Haocheng Zhong, Zhengyang Hu 0003 |
Softw. Pract. Exp. | 3 |
| 2025 | A Cross-Workload Power Prediction Method Based on Transfer Gaussian Process Regression in Cloud Data CentersabstractNowadays, machine learning (ML)-based power prediction models for servers have shown remarkable performance, leveraging large volumes of labeled data for training. However, collecting extensive labeled power data from servers in cloud data centers incurs substantial costs. Additionally, varying resource demands across different workloads (e.g., CPU-intensive, memory-intensive, and I/O-intensive) lead to significant differences in power consumption behaviors, known as domain shift. Consequently, power data collected from one type of workload cannot effectively train power prediction models for other workloads, limiting the exploration of the collected power data. To tackle these challenges, we proposeTGCP, a cross-workload power prediction method based on multi-source transfer Gaussian process regression.TGCPtransfers knowledge from abundant power data across multiple source workloads to a target workload with limited power data. Furthermore, Continuous normalizing flows adjust the posterior prediction distribution of Gaussian process, making it locally non-Gaussian, enhancingTGCP's ability to handle real-world power data distribution. This method enhances prediction accuracy for the target workload while reducing the expense of acquiring power data for real cloud data centers. Experimental results on a realistic power consumption dataset demonstrate thatTGCPsurpasses four traditional ML methods and three transfer learning methods in cross-workload power prediction. Ruichao Mo, Weiwei Lin 0001, Haocheng Zhong, Minxian Xu, Keqin Li 0001 |
IEEE Trans. Cloud Comput. | 3 |
| 2025 | TFEGRU: Time-Frequency Enhanced Gated Recurrent Unit With Attention for Cloud Workload PredictionabstractAccurate prediction of cloud workload is crucial for effective resource allocation in cloud computing. However, due to the complexity and high dimensionality of workloads in the cloud environment, achieving precise workload prediction is a complex and challenging problem. Current approaches to cloud workload prediction mainly rely on deep learning methods based on the Recurrent Neural Network (RNN), which struggle to capture the long-term dependencies inherent in workloads effectively. To tackle these challenges and overcome the limitations of existing methods, we propose an effective approach Time-Frequency Enhanced Gated Recurrent Unit with Attention (TFEGRU) for cloud workload prediction. First, we design a Time-Frequency Enhanced Block (TFEB) to capture complex workload patterns and extract features from both the frequency and temporal domains. Next, we integrate channel independent strategy and channel embedding into the model to adapt to high-dimensional workloads and enhance predictive performance. Finally, we apply a Gated Recurrent Unit (GRU) in conjunction with a multi-head self-attention mechanism to achieve accurate workload prediction. To validate the effectiveness of TFEGRU, comprehensive experiments are conducted using real-world traces from Google and Alibaba cloud data centers. The experimental results demonstrate that TFEGRU achieves accurate and efficient predictions across diverse cloud workloads, outperforming existing state-of-the-art methods. Feiyu Zhao, Weiwei Lin 0001, Shengsheng Lin, Haocheng Zhong, Keqin Li 0001 |
IEEE Trans. Serv. Comput. | 4 |
| 2024 | CycleNet: Enhancing Time Series Forecasting through Modeling Periodic PatternsabstractThe stable periodic patterns present in time series data serve as the foundation for conducting long-horizon forecasts. In this paper, we pioneer the exploration of explicitly modeling this periodicity to enhance the performance of models in long-term time series forecasting (LTSF) tasks. Specifically, we introduce the Residual Cycle Forecasting (RCF) technique, which utilizes learnable recurrent cycles to model the inherent periodic patterns within sequences, and then performs predictions on the residual components of the modeled cycles. Combining RCF with a Linear layer or a shallow MLP forms the simple yet powerful method proposed in this paper, called CycleNet. CycleNet achieves state-of-the-art prediction accuracy in multiple domains including electricity, weather, and energy, while offering significant efficiency advantages by reducing over 90% of the required parameter quantity. Furthermore, as a novel plug-and-play technique, the RCF can also significantly improve the prediction accuracy of existing models, including PatchTST and iTransformer. The source code is available at: https://github.com/ACAT-SCUT/CycleNet. Shengsheng Lin, Weiwei Lin 0001, Wentai Wu, Ruichao Mo, Haocheng Zhong |
NeurIPS | 6 |